Επιστροφή στις Ειδήσεις
ΚαινοτομίαAI Understanding ενημέρωση

Οι ερευνητές εισάγουν ένα συμπαγές μοντέλο θεμελίωσης για την επεξεργασία τσεχικών εγγράφων HTML

Οι ερευνητές εισήγαγαν ένα συμπαγές μοντέλο θεμελίωσης που ονομάζεται HTML-LM, σχεδιασμένο για την επεξεργασία τσεχικών εγγράφων HTML. Το μοντέλο έχει 154 εκατομμύρια παραμέτρους και εκπαιδεύτηκε σε 100 εκατομμύρια διαδικτυακά έγγραφα χρησιμοποιώντας πολλαπλούς στόχους.

4 min readRead the primary source
Source-page capture accompanying Researchers introduce compact foundation model for Czech HTML documents processing
Έγγραφο κύριας πηγήςΗ πηγή καταγράφηκε
Εκδότης
arxiv.org
Σύνδεσμος πηγής
arxiv.orghttps://arxiv.org/abs/2609.18494
Τύπος πηγής
Κύριο έγγραφο — μια επίσημη ανακοίνωση, χαρτί, αρχειοθέτηση ή σελίδα πρώτου μέρους που διαβάζουμε απευθείας.
ΠλαίσιοΚαταλάβετε αυτό σε 60 δευτερόλεπτα

Ξεκινήστε εδώ

Βασικοί όροι

Μοντέλο θεμελίωσης
Ένα μεγάλο προεκπαιδευμένο μοντέλο που μπορεί να προσαρμοστεί σε πολλές μεταγενέστερες εργασίες.
Ταξινόμηση
Μια εργασία όπου ένα μοντέλο εκχωρεί μια είσοδο σε μία ή περισσότερες προκαθορισμένες κατηγορίες.
Σύνολο δεδομένων
Μια συλλογή δομημένων ή μη παραδειγμάτων που χρησιμοποιούνται για εκπαίδευση, επικύρωση ή δοκιμή.
Δοκιμάστε τον εαυτό σαςΤι είναι το AI; Κουίζ

Τι έγινε

Researchers have developed a compact called HTML-LM, which is designed to process Czech HTML documents. The model has 154 million parameters and was trained on 100 million web documents using multiple objectives. It sets a new state-of-the-art for and regression applications in the Czech Internet domain.

The model has 154 million parameters and was trained on 100 million web documents using multiple objectives.

It sets a new state-of-the-art for and regression applications in the Czech Internet domain.

Στοιχεία πηγής: arxiv.org ↗

Γιατί έχει σημασία

The development of HTML-LM is significant because it addresses the limitations of existing approaches to processing web documents. It is a compact model that is both performant and economic, making it suitable for high-traffic industrial environments.

The model is trained on a large of web documents, which allows it to learn structural information inherent in HTML.

It uses a ModernBERT-based architecture, which enables it to process real-world web pages effectively.

The model is deployed in production, processing thousands of web documents per second, making it a practical solution for industrial environments.

It is released to the community under the CC BY-NC 4.0 license, making it available for use by others.

The model sets a new state-of-the-art for and regression applications in the Czech Internet domain, surpassing both larger encoders and small-sized LLMs.

Interactive Mechanism

Διαδραστικός Μηχανισμός: Πώς λειτουργεί στην πραγματικότητα

Εξερευνήστε την υποκείμενη τεχνολογία πίσω από αυτήν την εξέλιξη διαδραστικά.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Διαδραστικός Έλεγχος Έννοιας+10 Points
What is AI? Quiz

A route planner searches possible journeys using explicit rules. What does this illustrate about AI?

Τι να παρακολουθήσετε στη συνέχεια

The development of HTML-LM is a significant step towards creating universal, high-quality representations of web documents in high-traffic industrial environments.

The model's performance and economic viability make it a promising solution for industrial environments.

The model's ability to process real-world web pages effectively makes it a practical solution for a wide range of applications.

The model's release to the community under the CC BY-NC 4.0 license makes it available for use by others.

The model's deployment in production demonstrates its ability to process thousands of web documents per second.

The model's performance in the Czech Internet domain sets a new state-of-the-art, surpassing both larger encoders and small-sized LLMs.

Σχετικοί οδηγοί και κουίζ

Τι είναι το AI;Επεξήγηση μοντέλων AIΜετασχηματιστέςΕκπαίδευση AIΔοκιμάστε τι γνωρίζετε — δοκιμάστε ένα δωρεάν κουίζ AIΑναζητήστε έναν όρο AI στο γλωσσάρι μαςΑκολουθήστε τον ιχνηλάτη έκδοσης μοντέλου AI
Βρήκατε αυτό χρήσιμο;