GAIT: Graph-Memory Agent with Interrupt-Triggered Reasoning

Demo

GAIT (Graph-memory Agent with Interrupt-Triggered reasoning) builds a persistent, structured memory for vision-language-action (VLA) models by constructing semantic scene graphs online during navigation. The agent incrementally builds a graph-based spatial memory from egocentric observations, encoding object identities, spatial relations, and semantic attributes as it explores. An interrupt-triggered reasoning mechanism allows the agent to re-plan on the fly when novel or task-relevant cues are detected — bridging reactive control with deliberative graph-based reasoning.

This enables open-vocabulary navigation: given a free-form language goal, the agent retrieves and reasons over its scene graph memory to localize targets it has previously observed, or to plan exploration strategies for unseen goals — without requiring a fixed object vocabulary or pre-built map.

Status: More to come soon!

Erwin POUSSI
Erwin POUSSI
Aeronautics & Astronautics @ Stanford | Research Assistant (MSL & Qiu Lab)

Graduate student in Aeronautics & Astronautics at Stanford, working on robot learning and autonomous systems, with a focus on sim-to-real transfer and learned control policies.