Tilbake til Nyheter
InnovasjonAI Understanding orientering

Naturlig språkkodehenting for 1C:Enterprise: En åpen benchmark og effektiv bi-koder

Henting av naturlig språkkode er en oppgave i rask utvikling innen informatikk. Imidlertid kombinerer 1C:Enterprise-økosystemet russisk syntaks med svært domenespesifikk terminologi, for hvilke åpne datasett og spesialiserte modeller har vært praktisk talt ikke-eksisterende.

5 min readRead the primary source
Source-page capture accompanying Natural Language Code Retrieval for 1C:Enterprise: An Open Benchmark and Efficient Bi-Encoder
PrimærkildedokumentKilde registrert
Utgiver
arxiv.org
Kilde lenke
arxiv.orghttps://arxiv.org/abs/2608.19957
Kildetype
Primærdokument – en offisiell kunngjøring, papir, arkivering eller førstepartsside vi leser direkte.
KontekstForstå dette på 60 sekunder

Start her

Nøkkelord

Bi-koder
En modell som koder forespørsler og dokumenter i separate vektorer slik at de kan sammenlignes raskt i skala.
Benchmark
En standardisert test eller datasett som brukes til å måle og sammenligne modellytelse.
Henting
Finne relevante dokumenter eller poster fra en kunnskapskilde for en spørring.
Test deg selvHva er AI? Quiz

Hva skjedde

The authors present a comprehensive pipeline for 1C code : an open of 3,413 real-world, PII-scrubbed query-code pairs, a reproducible evaluation harness, and a specialized . To overcome scarce labeled data, they fine-tune on 784,057 synthetic triplets generated by google/gemma-4-26B-A4B-it from public code repositories, using Matryoshka Representation Learning (MRL) and a privacy-aware tokenizer.

The authors present a comprehensive pipeline for 1C code , consisting of an open , a reproducible evaluation harness, and a specialized .

The open consists of 3,413 real-world, PII-scrubbed query-code pairs, which are used to evaluate the performance of the proposed pipeline.

The reproducible evaluation harness ensures that the evaluation process is consistent and can be reproduced by others.

The specialized is designed to efficiently retrieve code snippets from the 1C:Enterprise ecosystem, which combines Russian syntax with highly domain-specific terminology.

To overcome scarce labeled data, the authors fine-tune on 784,057 synthetic triplets generated by google/gemma-4-26B-A4B-it from public code repositories, using Matryoshka Representation Learning (MRL) and a privacy-aware tokenizer.

The proposed pipeline's ability to efficiently retrieve code snippets from the 1C:Enterprise ecosystem is a significant advancement in the field of natural language code .

The use of the open and reproducible evaluation harness ensures that the evaluation process is consistent and can be reproduced by others.

The proposed pipeline's potential to improve the development of 1C code systems makes it an important area of research and development.

The proposed pipeline's ability to efficiently retrieve code snippets from the 1C:Enterprise ecosystem is a significant advancement in the field of natural language code .

The proposed pipeline's potential to improve the development of 1C code systems makes it an important area of research and development.

The proposed pipeline's ability to efficiently retrieve code snippets from the 1C:Enterprise ecosystem is a significant advancement in the field of natural language code .

The proposed pipeline's potential to improve the development of 1C code systems makes it an important area of research and development.

The proposed pipeline's ability to efficiently retrieve code snippets from the 1C:Enterprise ecosystem is a significant advancement in the field of natural language code .

The proposed pipeline's potential to improve the development of 1C code systems makes it an important area of research and development.

The proposed pipeline's ability to efficiently retrieve code snippets from the 1C:Enterprise ecosystem is a significant advancement in the field of natural language code .

The proposed pipeline's potential to improve the development of 1C code systems makes it an important area of research and development.

Kildedetaljer: arxiv.org

Hvorfor det betyr noe

The proposed pipeline addresses the lack of open datasets and specialized models for 1C code , enabling the development of more accurate and efficient retrieval systems. The use of MRL and a privacy-aware tokenizer also ensures the protection of sensitive information.

The proposed pipeline addresses the lack of open datasets and specialized models for 1C code , enabling the development of more accurate and efficient retrieval systems.

The use of MRL and a privacy-aware tokenizer ensures the protection of sensitive information, making the proposed pipeline more reliable and trustworthy.

The proposed pipeline has the potential to improve the development of 1C code systems, which is essential for various applications, such as software development and maintenance.

The use of synthetic triplets generated from public code repositories reduces the need for labeled data, making the proposed pipeline more efficient and cost-effective.

The proposed pipeline can be used as a starting point for further research and development of more advanced 1C code systems.

The proposed pipeline's ability to efficiently retrieve code snippets from the 1C:Enterprise ecosystem is a significant advancement in the field of natural language code .

The use of the open and reproducible evaluation harness ensures that the evaluation process is consistent and can be reproduced by others.

The proposed pipeline's potential to improve the development of 1C code systems makes it an important area of research and development.

The proposed pipeline's ability to efficiently retrieve code snippets from the 1C:Enterprise ecosystem is a significant advancement in the field of natural language code .

The proposed pipeline's potential to improve the development of 1C code systems makes it an important area of research and development.

The proposed pipeline's ability to efficiently retrieve code snippets from the 1C:Enterprise ecosystem is a significant advancement in the field of natural language code .

The proposed pipeline's potential to improve the development of 1C code systems makes it an important area of research and development.

Interactive Mechanism

Interaktiv mekanisme: Hvordan det faktisk fungerer

Utforsk den underliggende teknologien bak denne utviklingen interaktivt.

System Requirements:
Best ArchitecturePure RAGRecommended pattern
Hallucination RiskVery LowGrounding efficacy
Update Cost$0 (Vector sync)Ongoing maintenance
Core takeaway: Fine-tuning teaches models how to speak (form, style, syntax); RAG teaches models what to say (verifiable facts). Never use fine-tuning alone for factual memory.
Interaktiv konseptsjekk+10 Points
What is AI? Quiz

Which description best fits "narrow AI", the kind of AI in use today?

Hva du skal se neste

The authors' approach to fine-tuning on synthetic triplets generated from public code repositories, and the use of MRL and a privacy-aware tokenizer, are key aspects to watch in this research.

The authors' approach to fine-tuning on synthetic triplets generated from public code repositories is a key aspect to watch in this research.

The use of MRL and a privacy-aware tokenizer is another important aspect to watch, as it ensures the protection of sensitive information.

The proposed pipeline's ability to efficiently retrieve code snippets from the 1C:Enterprise ecosystem is also worth watching, as it has the potential to improve the development of 1C code systems.

The use of the open and reproducible evaluation harness ensures that the evaluation process is consistent and can be reproduced by others.

The proposed pipeline's potential to improve the development of 1C code systems makes it an important area of research and development.

The authors' use of synthetic triplets generated from public code repositories reduces the need for labeled data, making the proposed pipeline more efficient and cost-effective.

The proposed pipeline can be used as a starting point for further research and development of more advanced 1C code systems.

The proposed pipeline's ability to efficiently retrieve code snippets from the 1C:Enterprise ecosystem is a significant advancement in the field of natural language code .

The proposed pipeline's potential to improve the development of 1C code systems makes it an important area of research and development.

The proposed pipeline's ability to efficiently retrieve code snippets from the 1C:Enterprise ecosystem is a significant advancement in the field of natural language code .

The proposed pipeline's potential to improve the development of 1C code systems makes it an important area of research and development.

The proposed pipeline's ability to efficiently retrieve code snippets from the 1C:Enterprise ecosystem is a significant advancement in the field of natural language code .

The proposed pipeline's potential to improve the development of 1C code systems makes it an important area of research and development.

Relaterte guider og quizer

Hva er AI?ChatGPT og LLM-erKI-etikkAI-agenterAI-modeller forklartTransformatorerKIs fremtidAI treningPrompt EngineeringTest det du vet – prøv en gratis AI-quizSlå opp et AI-begrep i ordlisten vår
Fant du dette nyttig?