AI Papers Reader

Personalized digests of latest AI research

View on GitHub

NVIDIA Proposes Building AI Agents as Ordinary Python Objects

Building an AI agent often means assembling prompt templates, tool schemas, callback code, and workflow graphs, each in a different format. NVIDIA researchers argue that this split makes agents hard to learn, test, and maintain. In a preprint posted to arXiv in July 2026, they propose an alternative: write the agent as a plain Python class.

The framework, NVIDIA Object-Oriented Agents (NOOA), treats an agent as a Python object. Its methods are the actions a model can take, its fields hold durable state, its docstrings serve as prompts, and its type annotations act as input and output contracts. A method with an ordinary body runs as deterministic code. A method whose body is an ellipsis (...) is carried out at runtime by a language-model loop. The authors say this lets developers and models work through the same interface, so agent behavior can be tested and refactored like other software.

The default loop, called CodeAct, lets the model write and run Python in a session that holds the method’s arguments as live objects. Rather than pasting full data into the prompt, the harness shows a bounded preview. In one example from the paper, a list of 100 integers appears to the model as a short summary giving its type, its length, and its first and last five values. The full list stays in the execution environment, so the model can loop over all of it in code. The authors argue this lets agents work with data far larger than their context window. The loop ends only when the model returns a value that passes validation against the declared return type.

The authors tested the approach in several ways. On 88 targeted capability tests run five times across ten models, the suite passed 4,309 of 4,400 records (97.9%). Large and frontier models passed 99.2% of records, and small or efficient models passed 96.0%. Harder “stress” tests, involving batch bookkeeping, error recovery, and multi-step work, showed a wider gap: 93.9% for large models against 70.8% for small ones.

On agent benchmarks, the authors used a general-purpose agent of about 253 lines. On SWE-bench Verified, a set of 500 real GitHub repository issues, NOOA with GPT-5.5 at maximum reasoning effort reached 82.2%, compared with 78.6% for OpenCode and 78.2% for PI, two open coding harnesses. That is a 3.6-percentage-point lead over OpenCode. NOOA used about 1.1 million tokens per task, while PI used about 2.2 million. On Terminal-Bench 2.0, NOOA led with reasoning disabled, scoring 46.1% against 34.8% and 37.1%. At the highest setting, however, PI scored 75.3% to NOOA’s 73.0%.

The authors also report a CyberGym vulnerability-discovery result: 86.8% with GPT-5.5, which they describe as the top open-source entry in their comparison, though closed systems scored higher. On ARC-AGI-3, a benchmark of unfamiliar interactive grid games, a NOOA agent using GPT-5.5 scored 50.2% on the RHAE metric, compared with 41.7% for a baseline and 38.4% when its memory was replaced with plain files. With GPT-5.6-sol it scored 85.1%, at under $20 per game. The authors note that ARC Prize’s own evaluation of raw GPT-5.6-sol gave 13.3%, but that evaluation budgets differ.

Several caveats apply. The framework runs model-written code inside the agent’s own process, so its validation protects the agent loop but not the host machine; the authors acknowledge that sandboxing would give up the pass-by-reference behavior. The benchmark results come from the authors’ own harness and have not been peer reviewed. The comparisons are also sensitive to model and reasoning settings, and small models still struggle with multi-step tasks.

Still, the work offers a concrete argument that ordinary programming abstractions can serve as an agent interface without hurting model performance. If the approach holds up under independent testing, it could make agents easier for developers and coding assistants to build and inspect.