समाचार पर वापस जाएँ
नवीनताAI Understanding ब्रीफिंग

Google ने स्वायत्त एलएलएम प्रशिक्षण के बाद के लिए ऑटोफाइनट्यून की शुरुआत की

Google ने ऑटोफाइनट्यून जारी किया है, एक उपकरण जो टीपीयू पर एलएलएम पोस्ट-ट्रेनिंग प्रक्रियाओं के हाइपरपैरामीटर ट्यूनिंग और अनुकूलन को स्वचालित करने के लिए एआई एजेंटों का उपयोग करता है।

4 min readRead the primary source
Source-provided image accompanying Google introduces autofinetune for autonomous LLM post-training
प्राथमिक-स्रोत दस्तावेज़स्रोत रिकार्ड किया गया
प्रकाशक
developers.googleblog.com
स्रोत लिंक
developers.googleblog.comhttps://developers.googleblog.com/autonomous-llm-post-training-with-tunix-on-tpus/
स्रोत प्रकार
प्राथमिक दस्तावेज़ - एक आधिकारिक घोषणा, कागज, फाइलिंग, या प्रथम-पक्ष पृष्ठ जिसे हम सीधे पढ़ते हैं।
प्रसंगइसे 60 सेकंड में समझें

यहां से प्रारंभ करें

प्रमुख शर्तें

बड़े भाषा मॉडल (एलएलएम)
पाठ उत्पन्न करने और उसका विश्लेषण करने के लिए विशाल पाठ निगम पर प्रशिक्षित एक भाषा मॉडल।
प्रशिक्षण के बाद
प्रशिक्षण चरण पूर्व-प्रशिक्षण के बाद लागू किए जाते हैं, जैसे अनुदेश ट्यूनिंग, वरीयता अनुकूलन और सुरक्षा ट्यूनिंग।
लोरा (निम्न-रैंक अनुकूलन)
एक पैरामीटर-कुशल फ़ाइन-ट्यूनिंग विधि जो निम्न-रैंक एडाप्टर मैट्रिसेस जोड़ती है।
स्वयं की जांच करोएआई एजेंट प्रश्नोत्तरी

क्या हुआ?

Google released autofinetune, a system that automates LLM by using AI agents to iteratively optimize hyperparameters for Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL). The tool integrates Google's Tunix library, Gemma models, and Cloud TPUs, orchestrated via Antigravity CLI and Gemini Flash 3.7. In case studies, the agent autonomously adjusted parameters like LoRA ranks and learning rates, improving model accuracy and reward scores without manual intervention.

Google announced the release of autofinetune, a project designed to automate the of Large Language Models (LLMs). The system utilizes an autonomous research loop where an AI agent iteratively explores and optimizes training configurations. This approach is inspired by the earlier autoresearch project, which demonstrated autonomous pre-training exploration.

The tool leverages Google's full AI stack, specifically using the Tunix library for training, Gemma models as the base, and Cloud TPUs for compute. The orchestration is handled by Antigravity CLI and Gemini Flash 3.7. The primary goal is to replace the traditional manual cycle of adjusting hyperparameters with an automated process that runs experiments overnight and commits verified improvements to Git.

In the first case study, the agent optimized the google/functiongemma-270m-it model on the google/mobile-actions dataset using Supervised Fine-Tuning (SFT). The agent automatically adjusted parameters such as LoRA rank, alpha, optimizer, and learning rate. The results showed a consistent improvement in the model's ability to generate correct function calls, demonstrating the agent's ability to 'hill climb' toward better accuracy.

The second case study focused on Reinforcement Learning (RL) using the GRPO method to train Gemma 3 1B for math reasoning on the GSM8K dataset. RL is noted for its sensitivity to hyperparameters and instability. The autonomous agent identified better configurations for LoRA, rollout temperature, KL penalty, and system prompts. This resulted in an approximate 10% improvement in total reward, indicating better numerical and format accuracy in the model's answers.

स्रोत विवरण: developers.googleblog.com ↗

यह क्यों मायने रखता है?

This development significantly lowers the barrier to entry for high-quality LLM fine-tuning by removing the need for manual, repetitive experimentation. By automating the search for optimal hyperparameters, it allows developers to achieve better model performance with less specialized expertise and time. This shift toward autonomous research loops could accelerate the iteration cycle for AI developers, making advanced techniques more accessible and efficient for a broader range of organizations and individual researchers.

Autonomous addresses a significant bottleneck in AI development: the time and expertise required to manually tune hyperparameters. By automating this process, Google is making advanced model optimization more accessible to developers who may not have deep expertise in reinforcement learning or fine-tuning mechanics.

The integration of AI agents into the training loop represents a shift toward self-improving AI systems. If agents can reliably optimize their own training parameters, the pace of model improvement could accelerate, potentially reducing the cost and time associated with developing specialized LLMs for specific tasks.

This tool is particularly relevant for organizations using Google Cloud TPUs, as it provides a native, optimized workflow for leveraging this hardware. It also highlights the growing role of agentic AI in software engineering and research workflows, moving beyond simple code generation to complex experimental design and execution.

Interactive Mechanism

इंटरैक्टिव तंत्र: यह वास्तव में कैसे काम करता है

इस विकास के पीछे अंतर्निहित प्रौद्योगिकी का अंतःक्रियात्मक रूप से अन्वेषण करें।

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
इंटरएक्टिव कॉन्सेप्ट चेक+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

आगे क्या देखना है

Monitor the adoption of autofinetune in the developer community and any subsequent updates to the Tunix library. Watch for independent benchmarks that verify the performance gains claimed in Google's case studies, particularly regarding the stability of autonomous RL tuning. Additionally, observe if other major AI providers release similar autonomous tools, which could signal a broader industry shift toward self-optimizing model development pipelines.

Independent verification of the performance gains is crucial. While Google reports a ~10% reward improvement in the RL case study, third-party benchmarks will be needed to confirm these results across different datasets and model sizes.

The stability of autonomous RL tuning is a key area to monitor. RL is notoriously unstable, and it remains to be seen how well the agent handles edge cases or prevents reward hacking in more complex scenarios.

Adoption metrics for the autofinetune GitHub repository will indicate developer interest. If the tool gains significant traction, it may influence the broader ecosystem of LLM training libraries and tools.

संबंधित मार्गदर्शिकाएँ एवं प्रश्नोत्तरी

एआई एजेंटएआई प्रशिक्षणएआई मॉडल की व्याख्याआप जो जानते हैं उसका परीक्षण करें - निःशुल्क AI प्रश्नोत्तरी आज़माएँहमारी शब्दावली में एआई शब्द देखेंएआई मॉडल रिलीज ट्रैकर का पालन करें
क्या यह उपयोगी पाया गया?