GUIDE Sosiete

Kaaraange IA

AI security protects models, data, tools, and surrounding services from unauthorized access or manipulation.

2 simili jàngDañu mujjee yeesal Part of the AI Policy & Society learning path

Résumé

It includes ordinary software security and threats that target learning or model behavior. A secure design begins with the assets, adversaries, and trust boundaries of the actual application.

Takeaway yu am solo

  • Threat-model the full application.
  • Enforce permissions outside the model.
  • Retest controls across system changes.

Plongeur bu xóot

Identify what needs protection: private inputs, training data, model artifacts, credentials, connected accounts, and external actions. Record who can influence each input and what an attacker could gain from a failure. A public chatbot and an internal agent with write access have different threat models. Threats can affect different stages. Poisoned training material can alter learned behavior; adversarial inputs can manipulate predictions; untrusted retrieved content can redirect a tool-using application. Model output can also become dangerous when inserted into a database query, webpage, or command without appropriate handling. Apply controls at the software boundary. Enforce authorization in code, keep secrets out of model-visible context where possible, restrict tool scope, and validate outputs before use. A prompt asking a model to behave safely cannot replace account isolation or permission checks. Test representative failure paths in an authorized environment and maintain an incident process. Log enough information to investigate without collecting unnecessary sensitive content. Evaluate controls after changes to the model, retrieval sources, tools, and dependencies. Describe residual risk honestly; no single filter establishes complete protection.

Gis-gis xarala

A model refusing one malicious prompt does not prove that a system is secure. Different inputs, tools, modalities, and component boundaries can create distinct failure paths.

Locate the security boundary

  1. Imagine an assistant searching a private document store for a signed-in user.
  2. Apply the user’s access filter in the retrieval service before documents enter the model context.
  3. Test with a document belonging to a different account and verify that neither its contents nor identifying metadata appear in the result.

This defensive, hypothetical test checks authorization independently of the model’s willingness to follow instructions.

njeextalu pexe

Risk ak kaaraange

Gaañ-gaañu IA yu mag yi ak yu bës bu nekk yépp a ngi aju ci ki xam risk yi ak ki mëna def dara.

dogal yu gëna leer

Liggéeyukaay ak xam-xam bu ñépp bokk mooy wane ndax politiku kaaraange bu dëgër mën na am ci wàllu politik.

Cutting through hype

Faram-fàcce yu leer dañuy wàññi li ñuy jàpp ci hype, PR lab, ak tiyaatar bu leerul.

Doxal ci àdduna dëgg

Check that one account cannot retrieve another account’s documents.

Validate generated fields before using them in a database operation.

Risk yi ak balustrade yi

Jàppale risku nekk gi ni siyaas fiksioŋ fekk kàttan gi dafay yokk.

Jaxasoo kaaraange produit surface ak jubluwaay ci suufu autonomie bu kawe.

Bàyyi nit ñi xamul làkku Àngle ak ñi xamul làkku Angale, ñu am balluwaay yu baaxul.

Roadmap ngir samp gi

1

Tàqale loraange yi ci produit bi, jëfandikoo bu baaxul, ak risku ñàkka mëna yor / ñàkka méngoo.

2

Laajteel ban firnde mooy soppi sa xalaat ci kalendriye yi ak tar gi.

3

Danga taamu balluwaay yu njëkk yi ak jàngat yu fëgër yi moo gën waxtaanu njaay mi.

4

Xaarandil benn yoonu jëf: liggéey, politik, xaalis, wala xam-xam — du xam-xam kese.

Sources ak leneen luñu ci mëna jàng

Weyal di banneexu

Free newsletter

Get the daily AI briefing

Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the AI Security quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

Tambalil quiz

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

Next in AI Policy & Society

Kaaraange IA

Laaj yi ñuy faral di laaj

Is a strong system prompt enough to secure an assistant?

No. Authentication, authorization, input and output handling, tool limits, and incident response remain necessary parts of the application.