Kembali ke Berita
InovasiAI Understanding pengarahan

Para peneliti mengadaptasi perangkat Dolma Ai2 untuk membangun korpus pelatihan bahasa Thailand yang lebih baik

Sebuah tim peneliti Thailand mengadaptasi perangkat kurasi data terbuka Ai2 untuk membuat Mangosteen, sebuah korpus dengan 47 miliar token yang menurut tim meningkatkan kinerja model berbahasa Thailand sekaligus menghapus sebagian besar data web yang mendasarinya.

5 min readRead the primary source
Primary-source image accompanying Researchers adapt Ai2’s Dolma toolkit to build a better Thai training corpus
Dokumen sumber utamaSumber direkam
Penerbit
allenai.org
Tautan sumber
allenai.orghttps://allenai.org/blog/thai-llm-dolma
Jenis sumber
Dokumen primer — pengumuman resmi, makalah, pengarsipan, atau halaman pihak pertama yang kita baca langsung.
KonteksPahami ini dalam 60 detik

Mulai di sini

Istilah-istilah penting

Pra-pelatihan
Pelatihan model skala besar awal mengenai data luas sebelum adaptasi hilir.
Kekokohan
Kemampuan model untuk mempertahankan performa di bawah gangguan, pergeseran, atau masukan yang berlawanan.
Faktualitas
Seberapa akurat klaim model sesuai dengan informasi dunia nyata yang dapat diverifikasi.
Uji diri Anda sendiriKuis Penjelasan Model AI

Apa yang terjadi

According to Ai2, Thai researchers used its open Dolma toolkit to build Mangosteen, a 47-billion-token corpus for Thai language models. The team redesigned parts of Dolma to account for Thai text, data quality problems, and locally important sources that existing web-heavy datasets often missed.

Ai2 describes Mangosteen as a Thai corpus created by a team of Thai researchers who used Dolma as their starting point. The corpus contains 47 billion tokens. The team began with large web-data collections and sought to address weaknesses they saw in publicly available Thai datasets, including limited auditing by Thai speakers, unsuitable material, and insufficient coverage of books, research papers, official websites, and YouTube subtitles. Wannaphong Phatthiyaphaibun, identified by Ai2 as a project lead and a PhD student at the Vidyasirimedhi Institute of Science and Technology, said the team lacked enough developers to build a complete data-curation pipeline from scratch.

The researchers found that some Dolma components could not be transferred directly to Thai. In particular, they reported that sentence- and paragraph-level deduplication removed almost all of their data because Thai does not mark sentence boundaries in the same way as English. They retained document- and URL-level deduplication, modified other processing stages, adjusted quality filters, replaced language-specific tools, and added rules for patterns in Thai web data. One example involved Thai news pages that contained truncated snippets followed by “Read More” prompts; the team added filtering intended to remove those incomplete articles.

In experiments described by Ai2, the adapted pipeline removed more than 80% of the Common Crawl data used as a starting point and nearly half of FineWeb2, a collection the source describes as already cleaned and curated for model training. The team reported that models trained on the smaller, filtered datasets maintained or improved performance compared with models trained on the larger datasets. Ai2 also says those improvements carried into larger models and that Mangosteen-trained models performed better on Thai cultural-knowledge evaluations. The source does not give the evaluation names, numerical scores, model configurations, or experimental controls needed to assess the size and of those gains.

Detail sumber: allenai.org ↗

Mengapa itu penting

The work illustrates how language-model quality can depend on who curates the training data and whether the tools can be adapted to a language’s specific structure. The team reports that smaller, more carefully curated datasets matched or exceeded results from larger collections and improved performance on Thai cultural-knowledge evaluations.

The central significance is methodological: the project treats training-data curation as something that should be adapted by people who understand the language and its information environment. A general-purpose pipeline may preserve useful content, but its assumptions about sentence boundaries, duplication, quality, or web-page structure can behave differently across languages. The Thai team’s experience suggests that a pipeline designed around English-language conventions can discard valuable material or retain poor-quality material when applied elsewhere.

The reported results also challenge the simple idea that more training tokens automatically produce better language models. Ai2 says the team removed most of one web collection and almost half of another while maintaining or improving model performance. If independently reproduced, that would indicate that targeted filtering and source selection can sometimes deliver more value than adding raw text. Reducing irrelevant or incomplete material could also make training data easier to inspect, though the source does not report changes in training time, energy use, or total cost.

The reported improvement on Thai cultural-knowledge evaluations is particularly relevant to representation. A model can perform well on general language tasks while still missing local references, institutions, history, or cultural context. The team interprets its result as evidence that locally relevant data can help models represent the communities they serve. That conclusion remains a claim from the project’s experiments, not an independently established finding in this source. The article gives no detail about the benchmarks, the cultural domains tested, or whether gains came from source diversity, filtering, data volume, or other changes.

Interactive Mechanism

Mekanisme Interaktif: Cara Kerja Sebenarnya

Jelajahi teknologi yang mendasari di balik perkembangan ini secara interaktif.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Pemeriksaan Konsep Interaktif+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Apa yang harus ditonton selanjutnya

The source does not provide benchmark names, exact scores, model sizes, compute costs, corpus licensing details, or independent replication. The practical significance will depend on whether Mangosteen and the modified pipeline are released, whether other Thai researchers can reproduce the findings, and whether similar methods improve models in other underrepresented languages.

The first practical question is whether the research artifacts become available. The source emphasizes Dolma’s openness and describes Mangosteen as a reproducible adaptation, but it does not say whether the 47-billion-token corpus, filtering rules, replacement tools, or evaluation code are publicly released. Licensing and privacy considerations may affect what can be distributed, especially because the source says the project drew on web pages, books, research papers, official websites, and YouTube subtitles.

Independent evaluation will be important. The source reports comparisons with Common Crawl and FineWeb2, but it does not identify the models’ parameter counts, training procedures, data mixtures, benchmark scores, or statistical uncertainty. Follow-up work should test whether the gains persist across language generation, reasoning, , safety, and instruction-following tasks, rather than only the cultural-knowledge evaluations mentioned by Ai2. It should also examine whether filtering removed legitimate dialectal, regional, or minority-community material.

The broader test is whether this approach transfers to other languages and research teams with limited engineering resources. Dolma’s value, as presented here, is that researchers can modify an existing pipeline instead of rebuilding one. But the source does not establish that the same changes will work for languages with different writing systems, web ecosystems, or data availability. Future updates should clarify reproducibility, corpus governance, source coverage, model behavior, and the trade-offs between local relevance, data quality, openness, and responsible use.

Panduan & kuis terkait

Model AI DijelaskanPelatihan AIEtika AIUji pengetahuan Anda — coba kuis AI gratisCari istilah AI di glosarium kamiIkuti pelacak rilis model AI
Apakah ini berguna?