ወደ ዜና ተመለስ
ፈጠራAI Understanding አጭር መግለጫ

የጄቭ ውሳኔ ሞዴል በClaude Opus 5 እገዛ Pokémon Redን በሳምንት ውስጥ ያጠናቅቃል

ገንቢ አንድሪው ቦይድ የውሳኔ አሰጣጡን ሂደት ለማሻሻል Claude Opus 5ን ተጠቅሞ የእሱ የጄቭ ውሳኔ ሞዴል ከሰባት ቀናት ባነሰ ጊዜ ውስጥ ፖክሞን ቀይን በተሳካ ሁኔታ ማጠናቀቁን ዘግቧል።

4 min readRead the original reporting
Source-provided image accompanying Jev decision model completes Pokémon Red in under a week with Claude Opus 5 assistance
ሪፖርት ተደርጓልምንጭ ተመዝግቧል
አታሚ
tomshardware.com
ምንጭ አገናኝ
tomshardware.comhttps://www.tomshardware.com/tech-industry/artificial-intelligence/developer-says-jev-decision-model-beat-pokemon-red-in-under-a-week-non-llm-engine-succeeds-where-traditional-chatbots-stalled-for-months-but-claude-opus-5-coached-the-model-through-its-dead-ends
የምንጭ ዓይነት
በዜና ማሰራጫ ሪፖርት ማድረግ - የአንደኛ ወገን ሰነድ አይደለም።

በግል ማረጋገጥ ያልቻልነው ነገር: ይህ የይገባኛል ጥያቄ በተሰየመው መውጫ ምክንያት ነው። በአንደኛ ወገን ሰነድ ላይ አላረጋገጥነውም። (tomshardware.com)

አውድይህንን በ60 ሰከንድ ውስጥ ይረዱት።

እዚ ጀምር

ቁልፍ ቃላት

ትልቅ የቋንቋ ሞዴል (LLM)
ጽሑፍን ለማፍለቅ እና ለመተንተን በትልቅ ጽሑፍ ኮርፖራ ላይ የሰለጠነ የቋንቋ ሞዴል።
ትክክለኛነት
በትክክል ትክክል የሆኑ የተተነበዩ አወንታዊዎች መጠን።
AI ወኪል
ብዙውን ጊዜ መሳሪያዎችን እና ማህደረ ትውስታን በመጠቀም ግቡን ለማሳካት የሚከታተል ፣ የሚያመዛዝን እና እርምጃዎችን የሚወስድ የሶፍትዌር ስርዓት።
እራስህን ፈትን።AI ወኪሎች ጥያቄዎች

ምን ተፈጠረ

A developer named Andrew Boyd, founder of Standard Agents Inc., announced that his Jev decision model successfully completed the game Pokémon Red, reaching the Hall of Fame on September 23, 2026. Unlike traditional large language model (LLM) agents that have historically required months to navigate the game, Jev operates as a decision-based engine that selects actions from a predefined list. The process was not entirely autonomous; the model relied on Anthropic’s Claude Opus 5 to monitor game logs and dynamically adjust the available options and their phrasing, effectively coaching the agent through complex sequences.

According to a report by Tom's Hardware, the Jev model achieved victory in Pokémon Red in under one week. The project was developed by Andrew Boyd, who leads Standard Agents Inc.

The model functions as a decision engine, restricting its inputs to a specific list of choices rather than generating free-form text or commands. This constraint is credited with the model's speed compared to previous attempts by LLM-based agents.

Claude Opus 5 played a critical role by monitoring the game's internal logs. It provided real-time adjustments to the available choices and their wording, helping the Jev engine overcome obstacles that had previously stalled other AI agents.

The gameplay was documented via a livestream, which allowed viewers to observe the agent's progress through the game's terminal or browser interface.

የምንጭ ዝርዝሮች: tomshardware.com ↗

ለምን አስፈላጊ ነው።

This development highlights a shift in architecture, moving away from pure LLM-based navigation toward specialized decision engines that leverage LLMs for high-level guidance. By offloading the 'coaching' or strategic oversight to a model like Claude Opus 5 while keeping the core execution engine constrained to a list of choices, the system achieved significantly faster results than previous chatbot-based attempts. This suggests that hybrid architectures—where a powerful model provides context and strategy while a lightweight engine handles execution—may be more efficient for complex, state-based tasks than relying on a single, monolithic model to interpret and act simultaneously.

The success of the Jev model demonstrates the potential of 'System One' style architectures, where a specialized, constrained engine handles the heavy lifting of execution while a more sophisticated model provides strategic oversight.

By separating the decision-making logic from the strategic coaching, the developer bypassed the latency and reasoning errors often associated with using LLMs for direct, step-by-step game control.

This approach provides a practical blueprint for developers looking to build AI agents that are more reliable and faster than current chatbot-based implementations, particularly in environments where and adherence to a rule set are paramount.

Interactive Mechanism

በይነተገናኝ ሜካኒዝም፡ በትክክል እንዴት እንደሚሰራ

ከዚህ ልማት በስተጀርባ ያለውን ቴክኖሎጂ በይነተገናኝ ያስሱ።

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
በይነተገናኝ ጽንሰ-ሐሳብ ቼክ+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

ቀጥሎ ምን እንደሚታይ

The primary unknown is the scalability of this hybrid approach to environments outside of structured, rule-based games like Pokémon Red. While the use of Claude Opus 5 as a 'coach' proved effective in this instance, it remains to be seen how such systems perform in dynamic, real-world applications where the 'list of choices' is not easily defined or where the environment is less predictable. Observers should monitor whether Standard Agents Inc. or other developers can generalize this 'decision model' framework to more complex software environments or enterprise workflows.

Future updates from Standard Agents Inc. regarding the underlying architecture of the Jev model and its potential application in non-gaming environments.

Whether this hybrid 'coach-and-engine' model can be applied to more complex, non-deterministic tasks where the 'list of choices' is not as clearly defined as it is in a retro video game.

Potential industry interest in the platform developed by Boyd, as the company seeks to commercialize the technology used to build these agents.

ተዛማጅ መመሪያዎች እና ጥያቄዎች

AI ወኪሎችAI ሞዴሎች ተብራርተዋልChatGPT እና LLMsየሚያውቁትን ይሞክሩ - ነፃ የ AI ጥያቄዎችን ይሞክሩበእኛ የቃላት መፍቻ ውስጥ የ AI ቃልን ይፈልጉየ AI ሞዴል መልቀቂያ መከታተያ ይከተሉ
ይህ ጠቃሚ ሆኖ ተገኝቷል?