Pada si Iroyin
AtunseAI Understanding finifini

Koodu lati Ṣakoso awọn oluṣakoso Python ifaseyin fun awọn iṣẹ ṣiṣe akoko gidi

Iwe tuntun arXiv ṣafihan koodu si Iṣakoso, ọna ti o jẹ ki awọn awoṣe ede nla ṣe ina awọn olutona Python ti awọn aye wọn ti wa ni aifwy laisi eyikeyi akoko asiko LLM, ṣiṣe yiyan igbese yiyara ju PPO lori Atari ati awọn ipilẹ MuJoCo.

4 min readRead the primary source
Source-provided image accompanying Code to Control synthesizes reactive Python controllers for real‑time tasks
Iwe aṣẹ orisun akọkọOrisun ti o gbasilẹ
Olutẹwe
arxiv.org
Orisun ọna asopọ
arxiv.orghttps://arxiv.org/abs/2609.38733
Orisun iru
Iwe akọkọ - ikede osise, iwe, iforukọsilẹ, tabi oju-iwe ẹgbẹ akọkọ ti a ka taara.
AtokọLoye eyi ni iṣẹju 60

Bẹrẹ nibi

Awọn ofin bọtini

Awoṣe Ede nla (LLM)
Awoṣe ede ti a ṣe ikẹkọ lori titobi ọrọ corpora lati ṣe ipilẹṣẹ ati itupalẹ ọrọ.
Isọpọ
Bii awoṣe ṣe daradara lori tuntun, data ti a ko rii ni ita eto ikẹkọ.
Agbara
Agbara awoṣe lati ṣetọju iṣẹ ṣiṣe labẹ ariwo, awọn iyipada, tabi awọn igbewọle ọta.
Ṣe idanwo fun ara rẹAwọn awoṣe AI ti ṣalaye adanwo

Kini o ṣẹlẹ

Researchers released a pre‑print titled “Code to Control: Synthesizing Parameterized Reactive Controllers” (arXiv:2609.38733v1). The work proposes a two‑stage pipeline: an LLM drafts the structural skeleton of a Python controller, then a derivative‑free optimizer searches for the numeric parameters that make the controller work in a given environment. Once trained, the controller runs as ordinary Python code, requiring no LLM calls or planning at decision time. Experiments on a suite of Atari games, Flappy Bird, and MuJoCo locomotion tasks show the method outperforms prior planning‑based program synthesis approaches, matches deep reinforcement‑learning baselines while using fewer environment interactions, and transfers across substantial changes in dynamics.

The authors describe a pipeline where a large language model (LLM) is prompted to produce a Python function that defines the skeleton of a controller—its control flow, conditionals, and high‑level actions. The generated code contains placeholder parameters (e.g., gains, thresholds) that are left unspecified.

A derivative‑free optimizer (such as CMA‑ES) then interacts with the target environment, evaluating the controller’s performance and adjusting the numeric parameters to maximize reward. Because the controller is pure Python, each evaluation incurs only the cost of running the environment, not LLM inference.

After convergence, the controller can be executed directly, with decision latency limited to the Python runtime overhead. The authors benchmark this latency against a standard PPO policy and report faster per‑step decision times under their timing protocol.

Empirical results span 20 Atari titles, the Flappy Bird game, and several MuJoCo locomotion scenarios (e.g., Hopper, Walker2d). Code to Control matches or exceeds the scores of PPO baselines while requiring roughly 30‑50% fewer environment steps. It also outperforms prior program‑synthesis methods that rely on online planning.

The paper includes a transfer experiment where the same synthesized controller is evaluated after a substantial change in environment dynamics (e.g., altered gravity). The controller retains performance better than PPO, indicating that the learned parameters capture robust control strategies.

Awọn alaye orisun: arxiv.org ↗

Kini idi ti o ṣe pataki

The approach tackles a key bottleneck in LLM‑driven control: latency caused by repeated model inference or planning at each timestep. By compiling the policy into native Python, Code to Control enables real‑time decision making that can be faster than traditional PPO policies, opening the door to latency‑sensitive applications such as robotics, autonomous vehicles, and interactive gaming. Moreover, the method achieves comparable performance with far fewer environment interactions, suggesting a more sample‑efficient path to high‑quality controllers. If the technique scales, it could reduce the compute cost of training control policies and simplify deployment, because the resulting controller is just a script that runs on standard hardware without needing a large language model at runtime.

Latency is a critical factor for control systems; eliminating the need for LLM calls at each timestep removes a major source of delay, making the approach viable for real‑time embedded applications.

Sample efficiency reduces the amount of simulated or real‑world interaction needed, which can lower training costs and accelerate development cycles for new tasks.

The method decouples policy representation from the underlying model, allowing the final controller to be deployed on hardware that cannot host large language models, broadening the range of possible use cases.

By demonstrating competitive performance on standard benchmarks, the work provides a proof‑of‑concept that LLM‑generated program synthesis can move beyond toy examples toward practical control problems.

Interactive Mechanism

Ibaraẹnisọrọ Mechanism: Bii O Ṣe Nṣiṣẹ Lootọ

Ṣawari imọ-ẹrọ abẹlẹ lẹhin idagbasoke yii ni ibaraenisọrọ.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Ibanisọrọ Erongba Ṣayẹwo+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

Kini lati wo tókàn

Future work will need to address several open questions: (1) how the method scales to higher‑dimensional, real‑world robotics tasks; (2) whether the LLM‑generated structures generalize across domains without extensive re‑synthesis; (3) the of the derivative‑free search to noisy or sparse reward signals; and (4) the availability of open‑source code and reproducibility packages. Watch for follow‑up papers that benchmark the technique on physical robots, for community releases of the synthesis pipeline, and for industry pilots that integrate the generated controllers into embedded systems.

Scalability to high‑dimensional, continuous control problems such as manipulation or aerial robotics, where controller structures may become more complex.

across tasks: whether a single LLM‑generated template can be reused with minor parameter tuning for multiple related environments.

to noisy rewards: derivative‑free search can be sensitive to stochasticity; future studies will need to test stability under realistic sensor noise.

Open‑source release: the community’s ability to reproduce the results depends on the authors publishing their code, prompts, and optimizer settings.

Industry adoption: watch for announcements from robotics firms or game developers that integrate Code to Control‑generated policies into production pipelines.

Awọn itọsọna ti o jọmọ & awọn ibeere

Awọn awoṣe AI ti ṣalayeAwọn aṣoju AIAI IkẹkọỌjọ́ Iwájú AIṢe idanwo ohun ti o mọ — gbiyanju idanwo AI ọfẹ kanWa ọrọ AI kan ninu iwe-itumọ waTẹle olutọpa idasilẹ awoṣe AI
Ṣe eyi wulo?