Επιστροφή στις Ειδήσεις
ΚαινοτομίαAI Understanding ενημέρωση

Ερευνητές εισάγουν το Blindspot Benchmark για την ασφάλεια τεχνητής νοημοσύνης μεγάλου ορίζοντα

Εισήχθη ένα νέο σημείο αναφοράς για την αξιολόγηση της ασφάλειας των πρακτόρων τεχνητής νοημοσύνης μεγάλου ορίζοντα. Οι ερευνητές εισήγαγαν ένα νέο σημείο αναφοράς που ονομάζεται Blindspot για την αξιολόγηση της ασφάλειας των πρακτόρων τεχνητής νοημοσύνης μεγάλου ορίζοντα. Το Blindspot αξιολογεί πλήρεις τροχιές χρήστη-πράκτορα-περιβάλλοντος μέσω προσαρμοστικής αντίπαλης αλληλεπίδρασης…

4 min readRead the primary source
Source-provided image accompanying Researchers Introduce Blindspot Benchmark for Long-Horizon AI Safety
Έγγραφο κύριας πηγήςΗ πηγή καταγράφηκε
Εκδότης
arxiv.org
Σύνδεσμος πηγής
arxiv.orghttps://arxiv.org/abs/2609.16305
Τύπος πηγής
Κύριο έγγραφο — μια επίσημη ανακοίνωση, χαρτί, αρχειοθέτηση ή σελίδα πρώτου μέρους που διαβάζουμε απευθείας.
ΠλαίσιοΚαταλάβετε αυτό σε 60 δευτερόλεπτα

Ξεκινήστε εδώ

Βασικοί όροι

AI Ασφάλεια
Ένα πεδίο που επικεντρώνεται στη μείωση της επιβλαβούς συμπεριφοράς, των αστοχιών και των κινδύνων κακής χρήσης σε συστήματα τεχνητής νοημοσύνης.
Σημείο αναφοράς
Μια τυποποιημένη δοκιμή ή σύνολο δεδομένων που χρησιμοποιείται για τη μέτρηση και τη σύγκριση της απόδοσης του μοντέλου.
στιβαρότητα
Η ικανότητα ενός μοντέλου να διατηρεί την απόδοση υπό θόρυβο, αλλαγές ή αντίθετες εισόδους.
Δοκιμάστε τον εαυτό σαςΤι είναι το AI; Κουίζ

Τι έγινε

Researchers have introduced a new called Blindspot for evaluating the safety of long-horizon AI agents. Blindspot evaluates complete user-agent-environment trajectories through adaptive adversarial interaction, stateful tool execution, and execution-grounded adjudication.

Blindspot is a live-simulation framework that allows for the evaluation of AI agents in various scenarios and domains.

The contains 22 attack families and 35 scenarios across seven domains, yielding more than 2,500 long-horizon trajectories.

Each trajectory is assigned one of five outcomes: Safe Completion, Correct Refusal, Unsafe Completion, Over-Refusal, or Indeterminate.

Blindspot is extensible, allowing for the addition of new attacks, scenarios, tools, policies, domains, and agent configurations without redesigning the evaluation pipeline.

The researchers evaluated 13 proprietary and open-weight LLMs using eight metrics covering unsafe completion, appropriate refusal, benign utility, over-refusal, repeated-run , and post-refusal failure.

Στοιχεία πηγής: arxiv.org ↗

Γιατί έχει σημασία

The introduction of Blindspot is significant because it provides a more comprehensive evaluation of , taking into account the agent's behavior over multiple turns and interactions. This is particularly important for long-horizon AI agents that operate in complex environments.

The introduction of Blindspot is significant because it provides a more comprehensive evaluation of .

The takes into account the agent's behavior over multiple turns and interactions, which is particularly important for long-horizon AI agents.

Blindspot is a step towards improving the safety of long-horizon AI agents.

The development of Blindspot will likely lead to the creation of new AI models that are safer and more reliable.

The will also help to identify areas where AI agents are failing and provide insights for improving their safety.

Interactive Mechanism

Διαδραστικός Μηχανισμός: Πώς λειτουργεί στην πραγματικότητα

Εξερευνήστε την υποκείμενη τεχνολογία πίσω από αυτήν την εξέλιξη διαδραστικά.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Διαδραστικός Έλεγχος Έννοιας+10 Points
What is AI? Quiz

A route planner searches possible journeys using explicit rules. What does this illustrate about AI?

Τι να παρακολουθήσετε στη συνέχεια

The development of Blindspot is a step towards improving the safety of long-horizon AI agents. It will be interesting to see how the is used in the development of new AI models and how it affects the field of .

The development of Blindspot is a step towards improving the safety of long-horizon AI agents.

It will be interesting to see how the is used in the development of new AI models.

The impact of Blindspot on the field of will be significant.

The will likely lead to the creation of new AI models that are safer and more reliable.

The development of Blindspot will also help to identify areas where AI agents are failing and provide insights for improving their safety.

Σχετικοί οδηγοί και κουίζ

Τι είναι το AI;Ηθική του AIΠράκτορες AIΕπεξήγηση μοντέλων AIΜετασχηματιστέςΔοκιμάστε τι γνωρίζετε — δοκιμάστε ένα δωρεάν κουίζ AIΑναζητήστε έναν όρο AI στο γλωσσάρι μαςΑκολουθήστε τον ιχνηλάτη έκδοσης μοντέλου AI
Βρήκατε αυτό χρήσιμο;