뉴스로 돌아가기
혁신AI Understanding 브리핑

Paper는 저장된 비트를 변경하지 않고 양자화된 언어 모델에 대한 되돌릴 수 있는 업데이트를 위한 CellFill을 제안합니다.

arXiv 논문에서는 원본 양자화된 가중치를 정확하게 유지하면서 별도의 업데이트 파일을 통해 배포된 4비트 언어 모델에 지식을 추가하는 방법인 CellFill을 제시합니다. 저자는 높은 사실 주입 비율과 되돌릴 수 있는 업데이트를 보고하지만 이 작업은 동료 검토를 거치지 않은 사전 인쇄 상태로 남아 있습니다.

6 min readRead the primary source
Primary-source image accompanying Paper proposes CellFill for reversible updates to quantized language models without changing stored bits
기본 소스 문서녹음된 소스
출판사
arxiv.org
소스 링크
arxiv.orghttps://arxiv.org/abs/2608.20873
소스 유형
기본 문서 — 우리가 직접 읽는 공식 발표, 논문, 서류 또는 자사 페이지입니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

메모리(에이전트 메모리)
AI 에이전트는 연속성을 향상하기 위해 여러 단계 또는 세션에서 사용하는 저장된 컨텍스트입니다.
양자화
모델 가중치를 8비트 또는 4비트와 같은 낮은 정밀도 형식으로 변환합니다.
교정
모델의 신뢰도 점수가 실제 정확성 확률과 얼마나 일치하는지입니다.
자신을 테스트해 보세요AI 모델 설명 퀴즈

무슨 일이 일어났나요?

Researchers describe CellFill, a method for adding knowledge to deployed quantized language models without changing the stored integer codes or shared scales in the original model file. The update is kept in a separate file that the authors call a fill and can be withdrawn by subtraction.

The paper starts from a property of 4-bit . Each model weight is represented by an integer code and a shared scale, while the quantization process discards an interval between stored values. The authors define “in-cell learning” as writing new information into that unused interval. Their stated constraint is that re-quantizing the resulting served weights must reproduce the original integer codes and scales exactly. In practical terms, the base model file can remain unchanged at the bit level while a separate update modifies the model’s behavior when applied.

CellFill is the paper’s implementation of this idea. It uses a low-rank position inside each cell and is trained on the vendor’s own 4-bit release. The source says the experiments cover NF4, quantization-aware-training and GPTQ-style grids, as well as Qwen3 and Gemma models ranging from 1.7 billion to 31 billion parameters. The authors say the method checks its guarantee in the integer domain for every constrained weight, rather than relying only on aggregate numerical measurements.

The authors report that CellFill wrote 83% to 97% of a corpus of real facts that the tested models verifiably did not know, with zero reported constraint violations across as many as 2.4 × 10^10 constrained weights. They also report tests intended to show that the added knowledge can be used rather than merely recalled: injected drugs could be composed with injected ingredients up to the model’s reported two-hop ceiling, and a software library released after the model’s cutoff was used in code at 2.6 times chance. These are claims from the preprint’s own experiments, not independently established results.

For a retrieval comparison on PopQA’s long-tail questions, the source says the fill answered 82% to 90% of questions that the released model could not answer, using 75 tokens per question. The cited retrieval systems used 90 to 1,008 tokens per question. The paper also says several writers can share one model release and that sequential updates consume a constant fraction of the remaining room. The source points to archived result files behind its tables, but the supplied text does not describe their contents in enough detail to assess the full experimental protocol.

소스 세부정보: arxiv.org ↗

왜 중요한가요?

If independently validated, the approach could make some model updates easier to audit, distribute, certify and revoke because the original release remains bit-identical while new behavior is carried separately. The source reports promising results, but it does not establish production readiness or broad reliability.

The practical significance is version control. A deployed language model may be tied to benchmark reports, certifications, device fleets or other systems that assume a particular file. A conventional update can create a new model artifact and complicate comparisons with the certified release. CellFill’s proposed separation between the unchanged base release and a removable fill could allow an organization to distribute new knowledge without replacing the original quantized file. That separation could also make reversibility more explicit. According to the paper, the update is a separate file that can be withdrawn by subtraction. This creates a potentially useful distinction between the durable base model and an update layer that can be added, removed or assigned to different users. Such a design might help operators track which knowledge was added and when, although the source does not present a complete audit, access-control or governance system.

The reported efficiency comparison points to another possible benefit. The fill answered many previously missed PopQA questions with a much shorter token budget than the retrieval systems used in the comparison. If the result survives replication and applies to broader workloads, a compact update could sometimes be cheaper or simpler to serve than repeatedly retrieving long context. The source does not provide enough information to determine whether the comparison controls for indexing, retrieval quality, latency, memory use or total system cost.

The method also changes how model updates might be evaluated. Because the authors require exact preservation of the quantized codes and scales, they offer a concrete, machine-checkable condition for the base artifact. That could make some update claims easier to verify than ordinary fine-tuning claims. However, bit identity is not the same as behavioral safety or quality: a model can behave differently through a fill even when its stored base bits are unchanged. The source does not show whether the approach protects against hallucinations, unwanted associations, data leakage or harmful newly injected knowledge.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Document Size:128K tokens
Needle Placement Depth (Location in document):50% into text
Attention Context Buffer Map:
Target Fact (50%)
Equivalent Pages~320Standard book pages
Retrieval Accuracy99.9%Needle recall score
RAM / KV Cache5.1 GBMemory overhead
Prompt CachingActive~80% discount on reuse
Core takeaway: Million-token context windows allow querying whole codebases or legal archives in one prompt. However, KV cache memory scales with context length, making prompt caching crucial for real-time production.
대화형 개념 확인+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

다음에 무엇을 볼 것인가

The main questions are whether outside researchers can reproduce the reported results, whether fills preserve existing capabilities, how multiple updates interact, and how much storage and runtime overhead the method adds. The preprint also leaves open questions about security, misuse and performance beyond the tested settings.

Independent replication is the first key test. The paper is an arXiv preprint, submitted on 21 August 2026 and revised on 24 August 2026, rather than a peer-reviewed publication in the supplied record. Researchers will need to verify the integer-domain guarantee, the zero-violation result and the reported knowledge-injection rates across the named schemes and model families. The archived result files may help, but the source text alone does not establish their accessibility, completeness or independent verification.

Future evaluations should examine whether the fill damages capabilities the base model already had. The supplied abstract emphasizes facts the models did not know, composition tests and a post-cutoff software library, but it does not report broad regression testing, refusal behavior, , factuality outside the injected corpus or performance on unrelated tasks. It also does not explain how the authors determined that the models verifiably lacked each fact before training.

The paper’s claims about sequential updates and multiple writers require careful testing. A method that works for one fill may behave differently as available cell space is consumed, as updates overlap or as writers introduce conflicting facts. Important practical measurements include the size of each fill, training and application time, inference latency, memory use, the number of sustainable updates and the conditions under which subtraction fully restores the previous behavior.

Security and governance are also unresolved. A removable update file could be useful for controlled deployment, but it could also become an independent tampering target if an attacker can replace or alter it. The source does not describe authentication, authorization, provenance, privacy protections or defenses against malicious knowledge injection. Watch for follow-up work testing revocation in real serving systems, interactions with safety tuning and performance on models, domains and update types beyond the experiments reported here.

관련 가이드 및 퀴즈

AI 모델 설명ChatGPT와 LLMAI 트레이닝AI 윤리알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 모델 출시 추적기를 따르세요.
이것이 유용하다고 생각하시나요?