HƯỚNG DẪN AI về ngôn ngữ

AI thực hiện như thế nào trong bài kiểm tra thanh

Các mô hình ngôn ngữ lớn như GPT-4 đã đạt điểm trên mức vượt qua điển hình trên các phiên bản mô phỏng của Bài kiểm tra thanh thống nhất.

  • đọc 4 phút
  • Cập nhật lần cuối
Trên trang nàyđọc 4 phút
  1. Tổng quan
  2. Lặn sâu
  3. Tác động chiến lược
  4. The Future of How AI Performs on the Bar Exam
  5. Triển khai trong thế giới thực
  6. Rủi ro & lan can
  7. Lộ trình thực hiện
  8. Tiếp tục khám phá
  9. Câu hỏi thường gặp

Tổng quan

The best-known figure is about 298 out of 400, reported in 2023. These results show that models can handle the written, rules-heavy format of a licensing test. They measure test performance under particular grading and comparison choices, not the full judgment a practicing lawyer needs.

Lặn sâu

The Uniform Bar Exam (UBE) has three parts. The Multistate Bar Examination (MBE) has 200 multiple-choice questions and counts for half the score. The Multistate Essay Examination (MEE) has six essays. The Multistate Performance Test (MPT) has two tasks in which the examinee works from a closed file of facts and law to write a memo or brief. Scores run to 400, and each jurisdiction sets its own passing line, commonly between 260 and 270. In late 2022, Michael Bommarito and Daniel Martin Katz tested GPT-3.5 on MBE practice questions. It answered about half correctly: well above chance, but below passing. In March 2023, Katz, Bommarito, Shang Gao and Pablo Arredondo reported that GPT-4 scored about 297 on a full simulated UBE. OpenAI's GPT-4 technical report cited roughly 298 and a percentile near the 90th. The percentile drew the sharpest criticism. In a 2024 paper in the journal Artificial Intelligence and Law, Eric Martínez showed that the 90th-percentile figure was based on February test takers. The February group includes an unusually large share of people retaking the exam after failing. Compared with first-time July takers, his estimate fell to around the 60th percentile, and lower on the essays. He also noted that the essays were scored by the researchers against published sample answers, not by trained bar graders, which adds uncertainty. Contamination is another concern: past questions circulate online and may have been in the training data. A common misconception is that passing the bar means a model can practice law. The exam tests recall of rules and structured analysis within a closed set of facts. Practice also requires investigating facts, counseling clients, planning strategy, judging risk and taking professional responsibility. In Mata v. Avianca (2023), lawyers filed a brief containing citations that ChatGPT had invented. The case showed that exam-level fluency does not guarantee reliable legal research.

Tác động chiến lược

Tốc độ và tỷ lệ

Quy trình công việc ngôn ngữ có thể di chuyển nhanh hơn mà không làm mất tính nhất quán.

Truy cập và tiếp cận

Nó mở rộng quyền truy cập vào các ngôn ngữ và phong cách giao tiếp.

Quyết định rõ ràng hơn

Các nhóm có thể dành nhiều thời gian hơn để đánh giá trong khi quá trình tự động hóa xử lý sự lặp lại.

The Future of How AI Performs on the Bar Exam

Bar scores are becoming less useful as a benchmark. Top models already score near the ceiling on multiple-choice legal questions, and contamination is hard to rule out. Researchers have moved toward task collections such as LegalBench, a collaborative benchmark released in 2023 that breaks legal reasoning into many narrower tasks. They have also moved toward tests built on real work, such as research memos and contract review, graded by practicing lawyers. The NCBE's NextGen bar exam, planned to begin in 2026, gives more weight to integrated skills such as research and drafting. That may make comparisons with AI more informative. Whether models can pass tests is largely settled. The open question is how reliably they perform, and how they fail, on unfamiliar real matters.

Triển khai trong thế giới thực

OpenAI's March 2023 GPT-4 technical report listed a simulated Uniform Bar Exam score of about 298 out of 400 and a percentile near the 90th, a figure that was then widely repeated in headlines.

A law professor runs past MBE-style multiple-choice questions through a chatbot and compares its answers with her students'. The model is strong on black-letter rules but less reliable when the facts are deliberately ambiguous.

An evaluator checks whether a practice question was posted online before a model's training cutoff. A model that has already seen the question may be recalling an answer rather than reasoning to it.

A bar applicant uses an AI tutor to explain the MBE questions she missed. She checks each explanation against a commercial outline because the tutor sometimes presents a minority rule as the majority rule.

Rủi ro & lan can

  • Sự thật ảo giác có thể lặng lẽ đi vào báo cáo, luồng hỗ trợ hoặc kết quả nghiên cứu.

  • Sự nhạy cảm kịp thời có thể tạo ra kết quả không nhất quán đối với các yêu cầu tương tự.

  • Dữ liệu văn bản nhạy cảm có thể bị lộ nếu khả năng kiểm soát quyền truy cập yếu.

Lộ trình thực hiện

  1. Xác định định dạng đầu ra, âm thanh và tiêu chuẩn chất lượng trước khi triển khai.

  2. Phản hồi mặt đất với các nguồn đáng tin cậy bất cứ khi nào độ chính xác quan trọng.

  3. Duy trì điểm kiểm tra đánh giá của con người đối với các kết quả đầu ra có mức độ rủi ro cao.

  4. Theo dõi các kiểu lỗi và đào tạo lại các lời nhắc hoặc quy trình làm việc thường xuyên.

Tiếp tục khám phá

Free newsletter

Get the daily AI briefing

Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the How AI Performs on the Bar Exam quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

Bắt đầu bài kiểm tra

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

Câu hỏi thường gặp

What is How AI Performs on the Bar Exam?

Large language models such as GPT-4 have scored above typical passing lines on simulated versions of the Uniform Bar Exam. The best-known figure is about 298 out of 400, reported in 2023. These results show that models can handle the written, rules-heavy format of a licensing test. They measure test performance under particular grading and comparison choices, not the full judgment a practicing lawyer needs.

Which Uniform Bar Exam component has 200 multiple-choice questions and counts for half the total score?

The MBE is the 200-question multiple-choice section and makes up 50 percent of the UBE score. The MEE is six essays, and the MPT is two closed-file tasks.

What simulated Uniform Bar Exam score did OpenAI's GPT-4 technical report cite in 2023?

OpenAI reported roughly 298 out of 400. That is above the 260 to 270 passing range common across jurisdictions, and it was paired with a percentile near the 90th.

According to Eric Martínez's 2024 critique, why was GPT-4's 90th-percentile figure misleading?

The February group includes a large share of people retaking the exam after failing. Compared with first-time July takers, Martínez estimated GPT-4 fell to around the 60th percentile overall.

Who graded GPT-4's essay answers in the 2023 simulated bar exam, according to the critique described in the guide?

The essays were scored by the study's authors against published sample answers, not by trained bar graders. Martínez noted this adds uncertainty to the written-section scores.

What did Bommarito and Katz find when they tested GPT-3.5 on MBE practice questions in late 2022?

GPT-3.5 got roughly half the questions right. That is well above the 25 percent expected from guessing on four-option questions, but below a passing level.