Robots that think about the future
This is the long version of what I spent my PhD on, written the way I'd explain it to a friend rather than to a review committee. If you'd rather have the formal version, there's a research statement PDF at the end.
Think about the last time you emptied a dishwasher. You probably didn't put the plates in whichever cupboard happened to be closest. You put them where plates go — which is to say, you spent a little extra effort now so that the next person (often future you) wouldn't have to hunt. Nobody told you to do that. There was no explicit instruction in the task "empty the dishwasher" that said "and leave the kitchen in a state that makes tomorrow easier."
Robots, almost universally, do not do this. And that turns out to be a surprisingly deep problem.
The trouble with finishing the task
Most planners deployed today are myopic. They take the task in front of them, find a good way to accomplish it, and stop. By any reasonable metric they succeed: the task is done, the cost is low, the benchmark score is high. The catch is that every action a robot takes also rearranges the world. Objects move. Surfaces fill up. Paths get blocked. None of that shows up in the score for the task that was just completed, and all of it is waiting for the next task.
What I find interesting here is that this isn't a perception failure or a control failure. The robot sees fine and moves fine. It's a framing failure: we told it that intelligence means finishing the current task, so it optimized exactly that, and the damage accumulates somewhere we never bothered to measure. Run that loop for a week in a real home and you get cluttered counters, blocked aisles, and storage nobody can find anything in — each individual decision locally perfect.
Giving planners foresight instead of replacing them
The tempting response is to throw out classical planning and train some enormous end-to-end policy to do better. I don't think that's the right move, and this is probably my strongest methodological opinion. Classical planners have something genuinely valuable: when they return a plan, that plan is correct. You get guarantees. Thirty years of work went into making them reliable, and "it leaves a mess" is not a reason to discard the reliability.
So instead of replacing the planner, I give it foresight. The formulation I worked on through my PhD is anticipatory planning: augment the usual task objective with a learned estimate of what future tasks will cost from the state you're about to create. The planner still does the planning. It just now prefers, among the plans that accomplish the task, the ones that leave the world in better shape.
The part I didn't expect going in is that you don't have to specify any of the resulting behaviors. Decluttering, keeping things reachable, staging objects where they'll be needed — none of that is coded in anywhere. It all falls out of a planner that has some model of what might be asked next. That felt like evidence we'd formalized the right thing. This was the ICRA 2023 paper.
Then geometry showed up and made it harder
Symbolic planning is a comfortable place to work because the world is discrete: the mug is on the table or in the cabinet. Reality is not so tidy. Where in the cabinet you set the mug determines whether a bowl behind it can ever be retrieved, and that's a continuous decision with effectively infinite options.
Carrying anticipation into task and motion planning was the RA-L 2026 paper. I've chased the same idea through navigation among movable obstacles too, where a robot shoves a box to the nearest spot that works and slowly turns a warehouse aisle into an obstacle course. Different domain, identical pathology: moving the obstacle solves today's problem, and where you leave it decides how usable the space is tomorrow.
An accident I liked
Here's my favorite result, mostly because we weren't looking for it. The whole point of learning a future-cost model is to reason about long-lived deployments. But that model turns out to be useful inside a single planning problem, as a way of guessing which samples are worth trying first — a placement that keeps the world open tends to also be a placement the rest of the plan can work with. Used that way it cuts planning time by about a third and search effort by about half, with no change to what the planner optimizes.
I presented that one at the Search Algorithms for Robot Learning workshop at IROS 2026 in Pittsburgh — the write-up is here. A thing built for the long horizon quietly paying for itself in the short horizon is the kind of result that makes me think the abstraction is load-bearing rather than decorative.
Where I want to take this
Anticipatory planning needs two things the real world refuses to hand over: a distribution over future tasks, and a goal state. "Tidy up the kitchen" supplies neither.
Foundation models for the parts that need common sense. LLMs and VLMs are good at exactly what's missing — reading context, guessing what someone meant, hypothesizing what a kitchen will be asked for next. They're also unreliable planners, because a semantically sensible suggestion can be unreachable or unstable. So I don't want to plan with them. I want to use them to build the pieces anticipatory planning runs on — semantic world models, task priors, goal specifications — and keep decisions with a planner whose consequences are verified rather than assumed. That split is what I'm working on now.
Agents that learn from their own deployment. A robot in the same house for a month keeps meeting the same layouts and the same routines, and its own history is a record of which arrangements worked and which caused obstruction later. That's training data for the future-cost model, free, continuously. Since true task distributions are never written down anywhere, learning them from deployment is what moves this past simulation.
Shared spaces. Homes, hospitals, and warehouses have other robots and people in them, and the environment is how they implicitly coordinate. One agent's convenient staging spot is another's blocked path, which means the future tasks worth anticipating include everyone else's.
Benchmarks that let the mess accumulate. This one may matter most, and it's the least glamorous. Embodied AI benchmarks almost all reset the environment between tasks, so long-term consequences are structurally invisible — an agent can score perfectly while behaving in exactly the short-sighted way that ruins real deployments. I want benchmarks where state persists across a long stream of tasks and the metrics are cumulative: total cost over the stream, how much accessibility degrades, how often the robot obstructs itself, whether success drifts downward the longer it runs. You cannot fix what the evaluation refuses to measure.
The short version
I'd like embodied agents to be judged not on whether they finished, but on what they left behind. Every one of these threads — foundation models, learning from deployment, multi-agent coordination, better benchmarks — is in service of that one shift.
The formal version of all this, with citations, is in my research statement (PDF).