뉴스로 돌아가기
혁신AI Understanding 브리핑

연구에 따르면 LLM의 언어 간 추론 기능이 항상 상호 교환 가능한 것은 아닙니다.

새로운 사전 인쇄에서는 여러 언어로 동일한 수학 문제를 풀 때 다국어 언어 모델이 공유된 내부 기능을 사용하는지 여부를 조사합니다. 기능 간의 기하학적 유사성이 기능적 상호 교환 가능성으로 안정적으로 변환되지 않는다는 것을 발견했습니다.

5 min readRead the primary source
Primary-source image accompanying Study finds cross-language reasoning features in LLMs are not always interchangeable
기본 소스 문서녹음된 소스
출판사
arxiv.org
소스 링크
arxiv.orghttps://arxiv.org/abs/2608.23809
소스 유형
기본 문서 — 우리가 직접 읽는 공식 발표, 논문, 서류 또는 자사 페이지입니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

데이터세트
학습, 검증 또는 테스트에 사용되는 구조화된 또는 구조화되지 않은 예제 모음입니다.
인코더
입력을 잠재 표현으로 변환하는 모델의 구성 요소입니다.
특징
예측을 위해 모델에서 사용되는 입력 변수입니다.
자신을 테스트해 보세요AI 모델 설명 퀴즈

무슨 일이 일어났나요?

Researchers analyzed five language models from four model families using multilingual grade-school math problems in English, German, French, Spanish, Russian and Chinese. They used sparse autoencoders and -swapping experiments to study whether models rely on shared internal representations across languages.

The preprint by Igor Bogdanov and Changcheng Huang investigates whether language models solve equivalent mathematical problems through common internal features or through language-specific computations that merely produce similar answers. The researchers used the Multilingual Grade School Math and retained problems for which the models produced valid reasoning traces in all six tested languages: English, German, French, Spanish, Russian and Chinese. They replayed those traces through each model and recorded internal representations at multiple layers. The source describes the work as an examination of five models from four model families.

The researchers first used Centered Kernel Alignment, or CKA, to identify layers where representations were aligned across languages. At each selected layer, they trained two sparse autoencoders. One was a reconstruction-only baseline. The other, introduced in the paper, was called a Geometry-Invariant Sparse Autoencoder, or GI-SAE. GI-SAE added an Information Noise-Contrastive Estimation loss designed to make the produce similar activations for traces representing the same problem, even when those traces used different languages or token positions. This gave the researchers a way to identify features that appeared geometrically shared across languages.

The paper then tested whether those apparently shared features actually played interchangeable roles. During a model’s forward pass, the researchers swapped values between languages and measured the resulting output change using Kullback-Leibler divergence per feature. According to the source, GI-SAE produced higher CKA and Jaccard similarity at nearly every layer. However, the increase in geometric similarity did not consistently produce greater functional interchangeability. The reported pattern was model-specific: GI-SAE strengthened cross-language structure in Qwen, produced no functional benefit in Gemma, and had mixed, layer-dependent effects in Llama and Phi. The work was accepted as a poster at the ICML 2026 Workshop on Mechanistic Interpretability.

Taken together, the experimental design compared representation geometry with the results of direct interventions. The multilingual traces provided matched problem contexts across English, German, French, Spanish, Russian and Chinese, while the sparse autoencoders provided the feature spaces used for comparison. CKA and Jaccard similarity described geometric alignment, and Kullback-Leibler divergence per feature described output change after swapping. The comparison therefore asked whether the features that looked shared also behaved as interchangeable components. Its answer varied by model family and layer: the geometric effect was broadly stronger with GI-SAE, while functional effects were not consistently stronger. The result is reported for the five models from four model families in the workshop paper.

소스 세부정보: arxiv.org ↗

왜 중요한가요?

The findings suggest that multilingual reasoning may be partly shared inside a model, but the location and practical usefulness of that sharing depend on the model architecture. This complicates efforts to interpret, audit or improve multilingual AI systems using surface-level similarity measures.

The central implication is that similar-looking internal representations should not automatically be treated as evidence that a model uses the same computation across languages. In this study, a method specifically designed to amplify shared geometry generally increased measured similarity, but that similarity did not reliably predict what happened when features were functionally intervened on. For people evaluating multilingual AI, the distinction matters: a model can organize information similarly across languages without allowing the corresponding internal features to be substituted safely or usefully.

The findings also point to architecture as an important variable in multilingual reasoning. The source reports that cross-language sharing appeared at different depths in different models, and that the practical effect of GI-SAE varied across Qwen, Gemma, Llama and Phi. That makes broad claims about how all multilingual language models reason less reliable. Interpretability tools may need to be calibrated to individual model families and layers rather than applied as if internal representations had a common structure.

There is a practical research value in separating geometric alignment from functional interchangeability. -level interventions are often used to investigate what a model is doing internally, and the paper suggests that similarity metrics alone may give an incomplete picture of whether a representation has a causal or operational role. The source does not establish that the method improves model accuracy, translation quality, safety or deployment performance. It reports an interpretability result in a controlled mathematical-reasoning setting, so its direct effect on users remains unknown.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
대화형 개념 확인+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

다음에 무엇을 볼 것인가

The main open questions are whether the result holds beyond grade-school mathematics and the six tested languages, whether the reported patterns replicate across model versions, and whether shared features can improve reliability or safety in deployed multilingual systems.

Further scrutiny should focus on the exact models, checkpoints, layer locations, sizes and effect sizes used in the experiments. Those details are not included in the source text provided here. The source also does not say whether the authors released code, trained autoencoders or evaluation data, so independent replication will be important for determining how robust the reported model-specific pattern is.

A key test will be whether the result extends beyond the Multilingual Grade School Math . Follow-up work could examine harder mathematics, factual question answering, translation, coding or other reasoning tasks, as well as languages with different scripts and linguistic structures. The current source does not establish that the same cross-language patterns occur outside the six-language mathematical setting or beyond the particular models studied.

Researchers should also test whether swapping changes measurable behavior such as answer accuracy, reasoning validity or error patterns, rather than only output distributions summarized by KL divergence. If shared features can be linked to reliable behavioral effects, they might become useful for multilingual debugging or targeted interventions. If not, the study’s main contribution will remain a caution that internal geometric resemblance is not, by itself, proof of shared computation.

The paper’s workshop acceptance provides a venue for discussion, but it does not resolve the study’s broader limitations. Important unknowns include how sensitive the results are to the selected reasoning traces, the choice of sparse-autoencoder settings and the choice of intervention layers. A full peer-reviewed evaluation and tests across more models would help establish whether the reported architecture-dependent behavior is a general property of multilingual language models.

관련 가이드 및 퀴즈

AI 모델 설명트랜스포머AI 트레이닝AI란 무엇인가?알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 모델 출시 추적기를 따르세요.
이것이 유용하다고 생각하시나요?