Neural basis of compositional control
Assia Chericoni*, Justin M. Fine*, Taha S. Ismail, Gabriela Delgado Salazar, Melissa C. Franch, Elizabeth A. Mickiewicz, Ana G. Chavez, Eleonora Bartoli, Danika L. Paulo, Vaishnav Krishnan, Mohamed Hegazy, Alica M. Goldman, Lu Lin, Garrett P. Banks, Nisha Giridharan, Mohammed Hasen, Nicole R. Provenza, Andrew Watrous, Seng Bum Michael Yoo, Sameer A. Sheth**, and Benjamin Y. Hayden**
Article
Research briefing
Supplement
Github

In recent years, neuroscientists have gotten really interested in behavior. Advanced video tracking methods and AI have allowed us to follow the detailed movements of animals (including monkeys) and humans with high resolution. Then we can use sophisticated behavioral modeling to divide that behavior - in an unsupervised way - into discrete states, and do exciting neuroscience with it.
This is the ethogramming approach. That name comes from field biology, which uses similar approaches, although done laboriously, by hand, by humans. That is more or less how scientists have formalized behavior for 100+ years. Divide behavior into discrete categories - groom, fight, forage, and string them together, serially, like beads on a bracelet.
So, from a scientific perspective, the order of operations is observe, divide, label. It’s the bottom up approach to understanding animal behavior. This is a really useful approach and it’s told us a lot. Our lab has done it too.
But there’s something missing from the whole approach. It misses how behavior is generated, and as a result, the observe-divide-label approach misses something about how behavior works.
Animals and humans don’t really generate our behavior like beads on a string. Really, what we do, is we start with goals. We have things we want. Those come first. The goals may change over time, but it’s the goals at any moment that drive what we do. And the behavior is second; so it is downstream of our goals.
If the goals are at the top, then a top-down approach to understanding the neural basis of behavior would be one that seeks first to identify the goals, then infer how the goals drive the behavior.
This makes a huge difference. That’s because goals are not all-or-none. We might have two goals at the same time. We might have a hierarchy of goals. One goal may be waxing and the other goal may be waning. A bottom-up approach is blind to all of that because it forces behavior at any time into one category - it’s either groom or fight, or forage. Never a blend.
So in this paper, we argue that the next step in neuroethology is goal inference (here's a couple of papers we love doing something kinda similar). Latent goal inference is harder than unsupervised approaches, which are basically fancy clustering in the high dimensional space of behavior. As scientists, we are blind to goals, and they can be tricky to infer from behavior. Is that dog running after a prey or running away from a predator? Or maybe a mix? Even if our subjects are humans, they often don’t know their own motivations, and if they do, can’t alway express them, especially how much each is affecting them. So to understand behavior, we as scientists have to do latent goal inference.
In neuroeconomics, we are used to thinking of choice as an all-or-none process. But in the real world it seldom is. Typically, in the real world, we can try to have our cake and eat it too, even if carefully designed laboratory experiments try to preclude that possibility. But we think neuroscience needs to open that door again. We need to study choice as a continuous process, one that can involved blended strategies, not just discrete choices. Just as microeconomic theory is a good foundation for discrete choices, control theory is a good foundation for continuous choices.
This study is about a really simple task experiment that doesnt have any grooming, any fighting, or any foraging, except in the simplest sense. It’s as absolutely stripped down at it can possibly be. As simple as possible, but no simpler.
It’s the prey-pursuit task, basically, an ultra-simple version of chasing, the same game kids and dogs love so much. The player uses a joystick to move an avatar on a rectangle pen, and goes after a moving prey. The prey has a tiny amount of AI (for video game nerds, it’s called A*) to evade the pursuer. On some trials there’s a predator.
We work with epilepsy patients at Baylor St Lukes hospital, people with a large number of electrodes implanted in their brains as part of their diagnosis proceduce. The surgeons place the electrodes, and it depends on where the neurologists think the epilepsy is coming from, but typically, it involves electrodes in the hippocampus, anterior cingulate cortex, and orbitofrontal cortex.
We will sometimes ask for half an hour of their time if they want to play a simple video game kind of like a simpler version of Pac-man.

Patients generally enjoy playing this game. They can typically do it zero-shot, meaning they can do it straight away, no special training needed.
We record their behavior. Then we used a control theoretic approach to decompose their behavior and figure out their moment to moment strategy. The potential goals in the task basically come down to (1) pursue prey 1; (2) pursue prey 2; (3) avoid the predator.
Our control models allowed us to figure out which strategy the person was performing at every moment of the task - including if that strategy was a blend of component strategies. And it usually was.
This is much trickier than it sounds. You can’t just ask whether the joystick is pointed at a prey - that’s misleading a lot of the time, because the best way to capture a prey is often not to head directly at it; you often want to approach in a circular path. Often another prey crosses your path and you ignore it, etc.
Behavior, generally, makes sense - people choose the prey that has the highest value and is easiest to capture, and that maximizes these two things. People change their strategy (not all-or-none, but a big shift in the composition of goals) when an unpursued prey crosses their path, etc.


We can predict when people change their minds - or at least, when they have a major shift in the composition of their strategies.

Then we recorded brain activity from single neurons in three brain regions as people performed the task. These are three key regions of the brain that are famous for things like executive control (ACC), navigation and memory (hippocampus), and reward processing (OFC). We found that they play a coordinated role in implementing complex dynamic interactive control. Specifically,
• ACC neurons predict major changes in policy blends (changes of mind) suggesting they integrate information relevant to changes of mind and compute the need to change.
• Hippocampal neurons encode and update the latent policy state supporting early planning. That is, they keep track of where we are in the system, representing a flexible map that can be referred to by the meta-controller.
• OFC activity is consistent with an encoding of the current value structure of the task, rather than policy switching.
Together these results are consistent with a tripartite functional division in which hippocampus serves as a state-estimating controller, ACC serves as a meta-controller, and OFC provides a value context signal
