A system built around the task
Reflexion’s work on agentic robotics has been developing since 2025 around a practical question: how can an intelligent agent use a robot’s capabilities, retain useful context and adapt its actions without rebuilding the application around every new model?
That question connects the work we have shared on memory, hierarchical execution and skill development. It also explains why Reflexion OS is an operating layer rather than a single robot model. The application needs to keep working as its context changes and as better capabilities become available.
For a robot owner, the starting point is concrete: an existing machine, a task to automate, and a result to obtain. The architecture matters because it determines how much work can be reused when the task, the model or the equipment evolves.
Public work on memory and execution
At GOSIM Hangzhou in September 2025, our Founder & CTO, Charbel Dandjinou, presented work connecting context engineering with robot memory. The subject was continuity: carrying relevant information beyond a single request and relating that context to action. GOSIM Hangzhou, 14 September 2025
His November 2025 architecture paper addressed another part of the system. It separated high-level planning from callable skills and independent physical enforcement, making the boundary between a requested action and its execution explicit. TechRxiv preprint, 18 November 2025
In May 2026, Charbel presented Skill Foundry at GOSIM Paris. The demonstration explored an agent embodied in a simulated robot, observing its own behavior and iterating on how to acquire skills. It brought execution feedback into the development process rather than treating a completed rollout as the end of the work. GOSIM Paris, 6 May 2026
These contributions address different parts of the same direction: an agent needs context, capabilities and a way to understand what its actions actually produced. The value comes from how these parts work together across a task, not from placing a language model beside a controller.
The convergence we see
Several developments in the field now make these architectural questions especially visible. Gemini Robotics 1.5 combines high-level embodied reasoning with a vision-language-action model. Its September 2025 announcement describes planning and tool use alongside physical execution. Gemini Robotics 1.5, Google DeepMind
The comparison is specific: separating a reasoning process from specialized execution gives the system different places to represent intent, evaluate progress and act. Reflexion’s work concentrates on the operating layer around those interactions.
Execution becomes part of learning
The 2026 publications on ENPIRE and ASPIRE develop another part of the picture. ENPIRE describes an agent-managed loop for evaluating and improving robot policies. ASPIRE uses execution traces to diagnose failures, revise programs and build a library of reusable skills. Both make feedback from action central to what the agent does next. ENPIRE, NVIDIA GEAR (June 2026)ASPIRE, NVIDIA GEAR (June 2026)
Our Skill Foundry presentation in May preceded these June publications. The useful comparison is not that the systems are identical. It is that observation, diagnosis and capability development have become explicit system functions across multiple lines of work.
This is the convergence we see: reasoning, specialized execution and accumulated experience are being connected through increasingly agentic systems. It is an architectural direction that we had already been developing and presenting publicly.
What this means for Reflexion OS
Our focus is the application built around these capabilities. A customer should be able to start from a task and a robot, then develop a system that can use suitable skills, remember relevant context and respond to what happens during execution.
We design Reflexion OS to preserve task-level logic as supported models, tools and robot configurations evolve. Skill Foundry extends that work toward acquiring and improving capabilities. Our memory work addresses continuity across tasks and across physical and digital environments.
Physical action also needs an independent execution boundary. A model’s reasoning and a robot’s authority to move are different responsibilities. Our work on the safety runtime addresses this boundary as part of the system, alongside the capabilities that make the application useful.
For us, the next step is to apply this architecture to concrete tasks with robot owners, manufacturers and integrators. Existing equipment matters as much as new platforms. A pilot provides a defined task, a setup and an opportunity to evaluate what the system makes possible.