A behaviour rewarded every single time collapses quickly once the rewards stop. A behaviour rewarded unpredictably survives long stretches of nothing. The asymmetry is one of the most reliable findings in learning research.
What the reward pattern teaches
An animal learning a behaviour is also learning what to expect afterwards. Continuous reward builds a precise expectation, and a single failure is immediately detectable as a change.
Variable reward builds no such precision. A run of failures is indistinguishable from the normal pattern, so there is no clear point at which the animal concludes the rule has changed.
Persistence therefore comes from ambiguity. The behaviour continues because the absence of reward has never meant anything in particular.
How this is used deliberately
Early training uses continuous reward, because a new behaviour needs an unmistakable signal about which action produced the outcome.
Once the behaviour is reliable, the reward is thinned gradually and irregularly. Thinning too fast breaks the behaviour; thinning at a fixed ratio teaches the dog to count.
The endpoint is a behaviour maintained by occasional, unpredictable payoff, which is what allows a trained recall to survive a walk where nothing is carried in a pocket.
The same mechanism creates problem behaviours
Owners rarely apply variable reward to behaviours they want to eliminate on purpose, but they apply it constantly by accident.
A dog that barks at the door and is ignored nineteen times, then let out on the twentieth, has been placed on exactly the schedule that produces maximum persistence.
Begging at a table follows the same pattern. Occasional success under specific conditions builds a behaviour far harder to remove than one fed every time.
Why inconsistent households struggle
Multi-person households generate variable schedules without anyone intending it. One person holds a rule and another does not.
From the dog's position this is not inconsistency but a single environment in which the behaviour pays sometimes. The result is stronger than either person's approach alone would produce.
Agreement across a household is more valuable for problem behaviours than any individual technique, because it removes the schedule that sustains them.
The extinction problem it creates
Removing a variably rewarded behaviour is slow, and it gets worse before it improves. The dog first tries harder, longer and louder.
Anyone giving in during that escalation has rewarded the intensified version, which becomes the new baseline. This is the single most common way well-intentioned attempts backfire.
The practical implication is to decide in advance whether the behaviour can be tolerated through the escalation. If it cannot, the behaviour is better managed through changing the situation than through waiting it out.