Làkk AI GUIDE

Tekki làkku masin

Machine translation is automated conversion from a source language to a target language.

2 simili jàngDañu mujjee yeesal

Résumé

Neural systems often treat this as a sequence-to-sequence task: read one sequence and produce another. Building and evaluating such a system requires attention to data alignment and language-specific errors.

Takeaway yu am solo

  • Check aligned data and document-level splits.
  • Record evaluation configuration.
  • Inspect meaning-changing errors by language pair.

Plongeur bu xóot

Parallel training data pairs source passages with corresponding translations. Incorrect alignment, duplicated material, or mismatched language labels can teach the wrong relationship. Clean the pairs and keep documents together when splitting evaluation data to avoid near-duplicate leakage. Tokenization determines how text becomes model inputs. A tokenizer that handles one writing system efficiently may split another into many more units. Check the length limits in both languages and whether truncation removes the end of either the source or target passage. Automatic metrics make repeated experiments practical. BLEU compares patterns of word sequences with reference translations, but a metric is not a complete judgment of meaning, readability, or suitability for a domain. Evaluation settings and reference choices matter, so record them with the score. Use an error taxonomy alongside metrics: additions, omissions, changed numbers, inconsistent terminology, incorrect negation, and awkward phrasing. Assess each language pair and domain separately. An average across several well-resourced languages can conceal failures in a less-represented language or specialized document type.

Gis-gis xarala

A sentence may have several valid translations. Low surface overlap with one reference does not necessarily imply incorrect meaning, while high overlap can still conceal a critical changed word.

Compare usefulness with word overlap

  1. Imagine a reference “The package did not arrive.” Candidate A says “The parcel never arrived.” Candidate B says “The package did arrive.”
  2. Candidate A uses different words but preserves the main meaning. Candidate B resembles the reference while reversing the outcome.
  3. Record the negation error explicitly instead of choosing a translation by appearance or overlap alone.

The constructed example demonstrates why metric-based comparisons need semantic review.

njeextalu pexe

Gaawaay ak yaatuwaay

Liggéeyukaay yi ci làkk yi mën nañu gëna gaaw te duñu yàq deggoo gi.

Dugg ak yegg

Dafay yaatal jëfandikoo gi ci làkk yi ak ci anam yi ñuy jokkoo.

dogal yu gëna leer

Ekip yi mën nañu gëna yàgg ci àtte ci jamono ji otomatisation di liggéey ci baamtu.

Doxal ci àdduna dëgg

Evaluate a fixed test set with both a documented metric and bilingual error review.

Audit source-target pairs for mismatched dates, names, and sentence boundaries.

Risk yi ak balustrade yi

Lépp lu jaarul yoon mën na dugg ci rapoor yi, jàppale ci liggéey bi, wala ci njariñu gëstu bi.

Sensibilite bu gaaw mën na jur njariñ yu wuute ci laajte yu noonu mel.

Done yu am solo mën nañu feeñ sudee seytu jëfandikoo gi néew doole.

Roadmap ngir samp gi

1

Mandargal formaa génne gi, melokaan bi, ak standard kalite yi laata ngay dugal ko.

2

Tontu yu am solo ak balluwaay yu wóor saa yu dëggu bi di am solo.

3

Fexeel am barabu xool nit ñi ngir am njariñ yu am solo.

4

Toppal anami gacce yi ak di faral di tàggataat ay laaj wala def-liggéey.

Sources ak leneen luñu ci mëna jàng

Weyal di banneexu

Free newsletter

Get the daily AI briefing

Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the Machine Translation quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

Tambalil quiz

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

Laaj yi ñuy faral di laaj

Is BLEU a percentage of correctly translated sentences?

No. It is an automatic reference-based metric, not a direct count of sentences that a human would judge correct.