뉴스로 돌아가기
혁신AI Understanding 브리핑

ReToolSQL은 도구 사용을 통해 SQL을 확인하고 복구하기 위해 31B 모델을 교육합니다.

새로운 arXiv 사전 인쇄에서는 다중 회전 도구 사용 궤적에 대한 지도 추론 추적과 강화 학습을 결합하는 2단계 교육 방법인 ReToolSQL에 대해 설명합니다. 저자는 31B Gemma 4 명령 조정 모델을 사용하여 BIRD-SQL의 개발 벤치마크에서 74.32%의 단일 패스 실행 정확도를 보고했습니다.

5 min readRead the primary source
Source-provided image accompanying ReToolSQL trains a 31B model to verify and repair SQL through tool use
기본 소스 문서녹음된 소스
출판사
arxiv.org
소스 링크
arxiv.orghttps://arxiv.org/abs/2608.27796
소스 유형
기본 문서 — 우리가 직접 읽는 공식 발표, 논문, 서류 또는 자사 페이지입니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

도구 사용
검색, 계산기 또는 API와 같은 외부 도구를 호출하는 모델의 기능입니다.
강화 학습
에이전트가 장기적인 수익을 극대화하는 행동을 학습하는 보상 신호를 통한 교육입니다.
미세 조정
사전 훈련된 모델을 특정 작업에 맞게 조정하기 위해 도메인별 데이터에 대한 지속적인 훈련입니다.
자신을 테스트해 보세요AI 에이전트 퀴즈

무슨 일이 일어났나요?

Researchers introduced ReToolSQL, a two-stage training framework for text-to-SQL systems. It first uses supervised on rejection-sampled reasoning traces, then applies agentic reinforcement fine-tuning over multi-turn tool-use trajectories. The paper reports that the resulting model can verify queries, retrieve evidence, and repair faulty SQL using execution feedback.

The arXiv record identifies ReToolSQL as a computer-science artificial-intelligence paper submitted on August 28, 2026. Its authors are Pratik Kakkar, Chandra Dhir, Ravi Shankar, Pareekshit Reddy Gaddam, and Anup Shirgaonkar. The paper focuses on text-to-SQL: generating SQL queries from natural-language questions. The authors argue that many existing approaches model this as a single-turn generation problem, which limits the ability to recover when a query fails or produces an incorrect result. This framing connects the paper’s training procedure directly to the error-recovery problem it sets out to address.

The proposed framework has two stages. First, supervised starts the model with rejection-sampled reasoning traces produced by a privileged teacher. The paper says this stage expands the set of questions the model can solve, measured through pass-at-k coverage on difficult cases. Second, agentic reinforcement fine-tuning operates on multi-turn tool-use trajectories. In this stage, the model is trained to decide when to verify a query, what evidence to retrieve, and how to repair SQL after receiving execution feedback. The stages therefore cover both an initial learned strategy and later interaction with tools during the query process.

The reported experiments use Gemma 4 instruction-tuned at 31 billion parameters. The paper says reinforcement alone reached 73.66% execution accuracy on the BIRD-SQL development benchmark, increasing to 74.12% with self-consistency. Starting reinforcement fine-tuning from the supervised checkpoint produced the strongest reported result: 74.32% execution accuracy in a single pass and 74.77% with self-consistency. The authors state that this ranked first on the BIRD single-model development-set leaderboard at the time of writing. The source does not provide the leaderboard’s full comparison set or the benchmark’s question count in its abstract. These figures distinguish the reported training variants while leaving the broader comparison context unspecified.

소스 세부정보: arxiv.org ↗

왜 중요한가요?

The work addresses a practical weakness in text-to-SQL systems: treating query generation as a single-turn task can limit error recovery. The reported results suggest that combining supervised initialization with tool-using may improve execution accuracy within a single dense model, without additional human annotation beyond the benchmark.

Text-to-SQL systems are intended to let a user express a database question in ordinary language while an AI system constructs an executable query. In that setting, a syntactically valid query is not enough: the query must execute correctly against the relevant schema and data. ReToolSQL’s central contribution, as described by its authors, is to make verification and repair part of the model’s learned behavior rather than treating generation as a one-shot event. That distinction matters because successful generation and successful execution are related but not identical outcomes.

The reported results are notable because the method uses one dense 31B model rather than a multi-model system. The paper also says its composite rewards are anchored on execution correctness and that it requires no human annotation beyond the benchmark itself. If the results hold outside the reported development setting, this could offer researchers and organizations a way to improve query reliability through training design and execution feedback rather than relying only on larger models or manually labeled examples. The result is therefore relevant to both model-training choices and the practical design of text-to-SQL systems.

The evidence remains limited. The source is an arXiv preprint, and the headline results are reported on a development benchmark rather than a production deployment or an independently described operational evaluation. The paper’s abstract does not establish how the method performs on other database schemas, unfamiliar domains, noisy requests, changing data, or security-sensitive environments. It also does not report the computational cost, response latency, tool-call frequency, or failure modes associated with the reinforcement-learning process. The leaderboard position is explicitly time-bounded and should be treated as the authors’ reported claim. Those limitations define what can and cannot be concluded from the reported result.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
대화형 개념 확인+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

다음에 무엇을 볼 것인가

The findings still need replication beyond the BIRD-SQL development benchmark. Important unknowns include performance on unseen databases, real enterprise workloads, tool-use cost and latency, the contribution of each training stage, and how reliably the method handles permissions, sensitive data, and ambiguous requests.

The first verification priority is replication on held-out and external evaluations. Observers should look for results on unseen databases, different schema designs, and questions that do not resemble the benchmark’s training distribution. Such tests would help determine whether the reported gains reflect broad robustness or optimization for BIRD-SQL’s particular tasks and execution environment. A convincing evaluation would need to make that distinction visible across the relevant settings rather than relying on the single reported development result.

The paper presents supervised and reinforcement fine-tuning as complementary, so ablation results will matter. Future work should clarify how much of the improvement comes from the warm-start checkpoint, how much comes from multi-turn , and whether self-consistency provides a dependable benefit relative to its extra computation. Reproducibility will also depend on disclosure of the teacher traces, reward construction, tool setup, and evaluation protocol, none of which is detailed in the supplied arXiv abstract. These details would help separate the contribution of the individual components and the overall workflow.

Practical deployment raises questions that the source leaves open. A system that retrieves evidence or executes trial queries may need controls around database permissions, sensitive records, query costs, and actions that change data. It is also not clear how the model behaves when execution feedback is incomplete, misleading, or unavailable, or when a natural-language request is ambiguous. Before describing ReToolSQL as enterprise-ready, evaluators would need evidence on these operational conditions, along with error analysis and human-review requirements. Those safeguards are part of assessing the method’s practical reliability, beyond its benchmark execution accuracy.

관련 가이드 및 퀴즈

AI 에이전트AI 모델 설명AI 트레이닝Prompt Engineering알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 모델 출시 추적기를 따르세요.
이것이 유용하다고 생각하시나요?