A mouse has drunk its fill of sugar water. It should no longer want more. Then a lever appears. An animal calculating the present value of the outcome ought to stop pressing: no reward will be delivered during the test, and pre-feeding has made the usual reward less desirable. If the mouse keeps pressing anyway, the act may have passed from “I do this to obtain that” into a different mode—one cued by the learned situation rather than updated by the value of the result.
This quiet five-minute test contains the central difficulty of habit science. From outside the box, a lever press is a lever press. Inside the brain, however, one system can select it by forecasting a wanted outcome while another can retrieve it quickly from a repeated context. The same muscles can move for different computational reasons. Counting the movement alone does not reveal which system is driving.
A study published July 27 in Nature Communications by a Kyoto University Graduate School of Medicine team asked a further question: Is becoming habitual the same neural process as performing a habit a great deal? Their answer was no. And they went beyond watching activity. By exciting or suppressing circuits and using light to erase learning-related synaptic strengthening, they selectively changed one dimension while sparing the other.
“Is It a Habit?” and “How Much?” Are Different Questions
The researchers trained food-restricted male C57BL/6J mice to press a lever for 20-percent sucrose solution. For three days, every press earned a reward. Six days of variable-ratio training followed: on average, 10 presses were required during the first two days and 20 during the next four. More pressing still tended to produce more reward, preserving a meaningful link between action and outcome.
Then, without a new cue or a change of room, the rule silently changed for four days. Only the first press after an irregular 30-to-90-second interval, averaging 60 seconds, paid out. Extra presses during the wait did not bring the next sugar drop closer. This variable-interval schedule, a classic instrument of learning research, loosens the contingency between how much an animal acts and how much it receives. It favors habitual control.
| Stage or test | What the mouse experienced | What it revealed |
|---|---|---|
| Continuous reinforcement, 3 days | One sucrose reward for every lever press | The basic action–outcome association |
| Variable ratio, 6 days | A reward after about 10, then about 20, presses | Goal-directed control in which more action can earn more outcome |
| Variable interval, 4 days | Only the first press after an average 60-second wait earns reward | Habit-favoring training where extra presses do not improve the outcome |
| Outcome devaluation | Thirty minutes of access to sucrose or chow, then five unrewarded minutes with the lever | Continued pressing after sucrose satiety indicates outcome insensitivity |
| Omission test | Not pressing for 20 seconds produces reward; a press resets the clock | Whether pressing persists when withholding action is now advantageous |
The decisive observation came after interval training. The mice divided into two patterns. Both groups had become habitual by the outcome-devaluation criterion: they pressed even after being sated on sucrose. But one maintained a high press rate while the other reduced its pressing. The lower-output animals were not therefore “less habitual.” Their fewer presses were still governed without reference to the outcome’s present value.
Circuit One: The Gate From ACC to RSC
The first pathway runs from the anterior cingulate cortex, or ACC, to the retrosplenial cortex, or RSC. The ACC is part of the prefrontal cortex and is implicated in monitoring outcomes, conflict, errors and decisions. The RSC helps connect memory, context, space and organized behavior. In goal-directed mice, synaptic transmission along this projection showed long-term potentiation, or LTP. As control became habitual, that strength diminished.
Correlation was not the final claim. Chemogenetically suppressing ACC activity facilitated the transition to habitual behavior. Exciting it pushed trained habitual mice back toward action controlled by current outcome value. Crucially, those manipulations did not change the amount of pressing in the matching way.
In the authors’ model, a potentiated ACC→RSC connection helps keep a goal-directed strategy in control. Training-related weakening opens a gate through which control can shift toward habit. This does not mean that the ACC is a warehouse where a habit is stored. Brains are distributed systems, and the projection is one regulator of the transition. Nor is the weakening brain damage; it is experience-related synaptic plasticity.
Circuit Two: The Volume From LOFC to Striatum
The second pathway runs from the lateral orbitofrontal cortex, or LOFC, to the central striatum, or CS. Orbitofrontal cortex is known for updating the value of choices and outcomes. The striatum is a hub for action selection, movement and reward learning. Habitual mice that maintained high pressing retained LTP in the LOFC→CS projection. Those that lowered execution had weaker transmission.
Calcium imaging through a head-mounted miniscope showed that population activity in this pathway around reward—during the interval-trained stage—correlated with how much the animal executed. Activity at the instant of the lever press did not. The team then erased LTP selectively at LOFC synapses projecting to the central striatum. Pressing fell, yet the animals remained insensitive to outcome devaluation and retained their general motivation for sucrose.
LOFC→CS therefore looked less like a switch that decides whether a strategy is habitual and more like a circuit that sustains the level at which an established habit runs. The advance is not merely naming two brain regions. It is refusing to collapse two behavioral questions—“Does the outcome still govern the action?” and “How often is the action performed?”—and intervening in each selectively.
- Slice electrophysiology: Researchers compared AMPA- and NMDA-receptor-mediated currents as a projection-specific measure of synaptic strength.
- Calcium imaging: A miniscope recorded population activity in projection neurons while freely moving mice pressed and received rewards.
- Chemogenetics and CALI: DREADDs raised or lowered neural activity; chromophore-assisted light inactivation using cofilin-supernova erased learned LTP in a selected projection.
From William James’s Ruts to a Synaptic Map
The idea that habit leaves a physical trace in the nervous system is old. In The Principles of Psychology in 1890, William James described living creatures as “bundles of habits” and tied repetition to the plasticity of nervous tissue. The image was almost that of a rut worn into a road. What James did not possess was a way to measure individual synapses.
In 1898, Edward Thorndike placed hungry cats in “puzzle boxes” and timed their escape after they happened upon the right latch or loop. His law of effect—the principle that satisfying consequences strengthen preceding responses—helped lay the ground for operant learning. Ivan Pavlov moved from digestion to conditioned reflexes, showing how a formerly neutral signal could acquire the power to call forth a response. B. F. Skinner later used levers and reinforcement schedules to measure how the timing and probability of consequences reshape rates of action. Kyoto’s lever, sweet reward, variable ratio and variable interval belong to that experimental lineage.
On the neural side, Donald Hebb proposed in 1949 that connections strengthen when cells are repeatedly active together. In 1973, Tim Bliss and Terje Lømo described enduring strengthening of synaptic transmission after brief high-frequency stimulation in the rabbit hippocampus. LTP became a leading cellular model for learning and memory. The Kyoto paper reads that form of strength in two specific prefrontal projections.
From the 1990s, Ann Graybiel and colleagues showed that activity in the striatum changes as action sequences become familiar. Firing that spreads through a task early in learning clusters near its beginning and end after extensive training—a pattern called task bracketing, consistent with “chunking” a sequence. During the 2000s and 2010s, work increasingly mapped goal-directed and habitual control onto interacting corticostriatal systems. In 2013, researchers reported that stimulating a lateral orbitofrontal–striatal pathway could suppress repetitive behavior in a mouse model.
1890 — William James connects habit with the plasticity of the nervous system.
1898 — Thorndike’s puzzle boxes quantify trial-and-error learning by consequences.
1904 onward — Pavlov systematizes conditioned reflexes and cue–response learning.
1949 — Hebb proposes activity-dependent strengthening as a rule of learning.
1973 — Bliss and Lømo detail long-term potentiation.
1990s — Basal-ganglia recordings expose “chunking” of learned action sequences.
2010s — Causal studies manipulate corticostriatal pathways governing habitual and repetitive behavior.
2026 — The Kyoto team separates transition to habit from volume of execution.
This Study Does Not Say a Habit Takes 66 Days
Human field research has found that repeating eating, drinking or exercise behaviors in a stable context gradually raises self-reported automaticity toward a plateau. A much-cited 2010 study by Phillippa Lally and colleagues followed 96 people for 12 weeks. Its median estimate was 66 days, but individual modeled times ranged widely, from 18 to 254 days. That finding is a useful antidote to a rigid “21-day rule.” It is not the answer to Kyoto’s mouse experiment: the species, behavior, timescale and definition of habit are different.
The soundest everyday lesson is not to attempt amateur brain stimulation. It is to recognize that habit is not one number. A person may preserve the strategy while changing the amount: the morning walk continues but grows shorter; checking a notification remains cue-driven but happens less often. Tracking behavioral control and execution on two axes can be more informative than calling every change success or failure.
Psychological reviews describe habits as context–response associations strengthened by repetition in stable settings. That helps explain why rearranging cues, adding friction to an unwanted action and making a desired response easier can be more dependable than relying on willpower alone. But those are inferences from a broad literature. The Kyoto researchers did not test a consumer behavior-change program.
OCD and Addiction: A Bridge in View, Not Yet Crossed
The authors suggest that dissecting habit control may eventually illuminate pathological repetitive behavior in obsessive-compulsive disorder, addiction and eating disorders. There is reason to investigate. Human and animal studies have associated orbitofrontal–striatal activity and connectivity with compulsivity. In mice, repeated stimulation of corticostriatal circuitry has produced persistent OCD-like behavior, while stimulation in another protocol suppressed repetitive actions.
But OCD is not a bad habit with a clinical label. It can involve intrusive thoughts, fear, ritual, distress and severe impairment, and its circuitry extends beyond one projection. A sated mouse pressing a lever and a person repeatedly checking a lock are separated by diagnosis, language, social meaning and subjective suffering. This paper locates possible bridge supports; it does not construct a treatment people can cross.
The future promise lies in refusing to label every symptom simply “repetition.” In one person, difficulty updating an outcome may dominate; in another, amplification of execution may matter more. If humans show a comparable separation, assessment and target selection could become more precise. For now, that remains a hypothesis to test.
A Strong Study Still Has Specific Limits
First, the subjects were male mice, 6 to 24 weeks old. Sex, age and species differences remain open. Human brains contain corresponding regions, but anatomical similarity does not prove an identical division of labor. Replication in female and older animals, other strains, different rewards and richer behaviors is needed.
Second, image-alignment limits meant the calcium-imaging experiments did not track the same individual neurons across training days. They compare populations at different stages. The study did not film one cell “changing jobs” from goal-directed to habitual control.
Third, the chemogenetic ligand CNO can convert back to clozapine in the body, creating a recognized risk of off-target effects. The authors used controls and converging CALI experiments based on a different mechanism. Agreement among methods strengthens the case without making any one method flawless.
Fourth, sample size differed by experiment. The principal behavioral training included 79 variable-interval mice and 19 controls, but devaluation, electrophysiology, imaging and manipulation used distinct subsets. The headline does not describe 79 mice each undergoing every technique.
- Does the circuit division hold in female, older and genetically different mice?
- Does it generalize beyond food to avoidance and longer action sequences?
- Can the same neurons be followed across days to establish the order of circuit change?
- Can human imaging or stimulation separate strategy from execution quantitatively?
- How do the two axes fail in pathological repetition, and how do existing therapies change them?
Freedom Is Not the Absence of Habit
Habit is often narrated as a defeat of will. Yet a brain that recalculated every consequence each time it brushed teeth, typed a familiar word or turned down a known street would not finish the day. Habits conserve attention, bind skilled movements into packages and free thought for novelty. As James understood, people are constrained by habits and liberated by them.
The danger is not automaticity by itself. It is failing to update when the world changes, or failing to lower execution when the act no longer serves. The Kyoto study shows that those can be different failures. The brain has a pathway that regulates which strategy controls behavior and another that sustains how strongly the behavior runs.
Behind the brief image of a sated mouse pressing a lever is a question more than a century old: Why do we continue an act after its reason has gone? And is continuing the same as doing too much? The new answer is limited to mouse circuitry. Even so, turning habit from one black box into two manipulable dimensions is a substantial conceptual step. Freedom may not mean erasing our habits. It may mean being able to close the gate—or turn down the volume—when purpose changes.
Reporting Notes and Principal Sources
This article is based on Kyoto University and Japan Science and Technology Agency materials, the original paper, and major literature on habit and synaptic plasticity available by July 29, 2026, 11:30 a.m. JST. “Gate” and “volume knob” are explanatory metaphors used by this article, not anatomical terms in the paper. Clinical discussion concerns research possibilities, not medical advice.
- Kyoto University: Neural pathways reveal why habits vary across individuals
- Kyoto University: detailed study summary, methods glossary and funding (PDF, Japanese)
- Nature Communications: Dissociable roles of prefrontal plasticity in decision-making strategy and execution of habitual behavior
- Japan Science and Technology Agency: research announcement (Japanese)
- Wood & Rünger: Psychology of Habit (2016)
- Lally et al.: How are habits formed? Modelling habit formation in the real world (2010)
- William James: The Principles of Psychology, Chapter IV, Habit (1890)
- Nobel Prize: Ivan Pavlov’s Nobel Lecture
- Chance: Thorndike’s puzzle boxes and the origins of experimental behavior analysis (PDF)
- Mitchell-Heggs et al.: Reflecting on 50 years of long-term potentiation
- MIT McGovern Institute: Striatal task bracketing and habit formation
- Burguière et al.: Orbitofronto-striatal stimulation suppresses compulsive behavior (Science, 2013)
