Tiếp theoHướng dẫn tiếp theo
Điểm chuẩn của đại lý mã hóa và băng ghế dự bị SWE
kỹ thuật
HƯỚNG DẪN KỸ THUẬT
Evaluating AI agents means measuring how well a multi-step system reaches goals in realistic settings.
That covers whether it finished the task, whether the path it took was sensible, whether it used tools correctly, what it cost, and whether it stayed safe. Agents act over many steps and have side effects, so checking a single answer against a reference is not enough, and good evaluations usually run the agent inside a controlled environment and inspect the final state.
Agent evaluation looks at several things. Task success asks whether the goal was actually achieved, ideally checked by looking at the environment afterward: tests pass, the right record exists, the file has the right contents. Trajectory quality asks whether the steps were reasonable, without loops, pointless calls or lucky guesses. Tool-use correctness checks whether the agent picked the right tools with valid arguments and handled errors. Cost and latency track tokens, calls and wall-clock time. Safety checks whether the agent avoided harmful or unauthorized actions and resisted prompt injection. Environment-based benchmarks have become a leading approach. SWE-bench, from Princeton researchers, asks agents to resolve real GitHub issues and checks the results with the project's tests, and SWE-bench Verified is a human-checked subset. WebArena provides self-hosted websites for browsing tasks. OSWorld tests agents operating full computer desktops. GAIA poses questions that need multi-step research and tools. The tau-bench suite simulates users and domain policies, and it introduced a pass^k metric that asks whether an agent succeeds on every one of k repeated attempts. Grading uses three main kinds of checks: code-based checks on outcomes, which are precise and cheap; model-based graders that score transcripts against a rubric, which are flexible but need to be checked against human judgment; and human review, which is the most trustworthy and the slowest. A common mistake is to trust public leaderboards alone. Benchmarks can leak into training data, become saturated, or fail to resemble your own workload. Another mistake is treating one run as the truth, since agents are nondeterministic and small differences in success rate can be noise. Teams get the most value from a private evaluation set built from their own real tasks and failures, run repeatedly and tracked over time.
Các quyết định về kiến trúc sẽ thúc đẩy hiệu suất và chi phí vận hành trong nhiều năm.
Giáo dục kỹ thuật giúp các nhóm chọn nhóm phù hợp chứ không chỉ nhóm mới nhất.
Lựa chọn kỹ thuật tốt hơn làm giảm sự cố về độ tin cậy trong sản xuất.
As agents take on longer tasks, evaluations are moving toward longer horizons, more realistic environments, and measures of reliability and cost as well as peak capability. Benchmark saturation and contamination will probably keep pushing developers to create new tasks and to rely more on private, domain-specific test sets. Safety and security evaluations, including resistance to prompt injection, are receiving more attention as agents get access to real systems. There is no agreed standard for grading open-ended agent behavior, so combining outcome checks, calibrated model graders and human review is likely to remain common practice.
A coding-agent team scores each attempt by whether the repository's hidden tests pass after the agent's patch, as SWE-bench does, instead of judging whether the diff looks right.
A customer-service agent runs against simulated users and a mock booking database. Graders check whether the final database state matches the correct outcome and whether the agent followed refund policy.
An engineering team runs every task five times and reports how often the agent succeeds on all five, which exposes flaky behavior that a single-run success rate would hide.
A company adds a safety suite in which tasks include a planted instruction to delete files or leak data, and it counts an otherwise successful run as failed if the agent obeyed.
Tối ưu hóa một điểm chuẩn có thể che giấu những điểm yếu của hệ thống rộng hơn.
Chi phí cơ sở hạ tầng và bảo trì thường được đánh giá thấp.
Khoảng cách về bảo mật và khả năng quan sát có thể tăng lên khi hệ thống trở nên phức tạp hơn.
Xác định các mục tiêu về độ trễ, chất lượng và chi phí trước khi triển khai.
Điểm chuẩn trong điều kiện tải và dữ liệu thực tế.
Giám sát thiết bị về lỗi, độ lệch và tác động của người dùng.
Chuẩn bị đường dẫn khôi phục và ứng phó sự cố trước khi mở rộng quy mô.
Free newsletter
Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.
One email each weekday. Unsubscribe in one click. We never sell or share your address.
Test yourself
Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.
Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation
Evaluating AI agents means measuring how well a multi-step system reaches goals in realistic settings. That covers whether it finished the task, whether the path it took was sensible, whether it used tools correctly, what it cost, and whether it stayed safe. Agents act over many steps and have side effects, so checking a single answer against a reference is not enough, and good evaluations usually run the agent inside a controlled environment and inspect the final state.
Multi-step actions change environments, so you need to check what actually happened and how, not just the final text.
Environment-based checks, here the project's tests, judge whether the fix really works.
Requiring success on all k attempts measures consistency, which matters for agents deployed to real users.
Trajectory checks judge the path taken, catching loops and waste that outcome checks alone miss.
Model graders are flexible but should be checked against human judgment and watched for systematic bias.
Tiếp tục học hỏi
Đã chọn thêm hướng dẫn cho chủ đề này
Tiếp theoHướng dẫn tiếp theo
Điểm chuẩn của đại lý mã hóa và băng ghế dự bị SWE
kỹ thuật