業界ガイド
AI in Courts and Judicial Risk Assessment
Judicial risk assessment tools are statistical models that estimate how likely a defendant is to miss court or be rearrested.
このページでは4 分で読めます
概要
Judges use the scores in bail, sentencing and parole decisions. The tools can make decisions more consistent and reduce unnecessary detention, but they raise hard questions about fairness, transparency and due process, as the debate over the COMPAS tool showed.
ディープダイブ
Risk assessment in criminal justice is older than machine learning. Many tools are fairly simple: a points system or regression model built from historical data. They score factors such as age, prior convictions, prior failures to appear and pending charges. The Public Safety Assessment (PSA), developed with funding from Arnold Ventures, draws on age and criminal history without interviewing the defendant. It produces separate scores for failure to appear and for new criminal activity, plus a flag for new violent criminal activity. COMPAS, sold by Northpointe (now Equivant), is a proprietary tool used in pretrial, sentencing and corrections decisions in some places. It also uses questionnaire answers. In 2016 ProPublica analyzed COMPAS scores in Broward County, Florida. It reported that Black defendants who did not reoffend were nearly twice as likely as white defendants to have been labeled higher risk, and that white defendants who did reoffend were more often labeled lower risk. Northpointe replied that the scores were equally predictive for both groups: a given score matched similar reoffense rates regardless of race. Both claims can be true at once. Researchers including Jon Kleinberg and Alexandra Chouldechova showed that when base rates differ between groups, a score generally cannot be calibrated and also have equal false positive and false negative rates. Choosing a definition of fairness is a policy decision, not a purely technical one. There are other concerns. Rearrest stands in for reoffending, and it reflects where police patrol. Proprietary models make scores harder to challenge. Judges may defer to scores, or override them selectively. Supporters answer that unaided judicial intuition is also biased and less consistent. A common misconception is that these tools are advanced AI. Most are transparent statistical models using a handful of factors. Most other AI in courts is administrative: sorting e-filings, transcription, translation and scheduling.
戦略的影響
背景とルール
AI のアイデアが現実と接触しても生き残れるかどうかは、業界の状況によって決まります。
品質管理
ドメインの制約は、許容可能なエラー率と監視モデルに影響を与えます。
ビルドの選択
導入を成功させると、技術的能力と最前線のワークフローが連携します。
The Future of AI in Courts and Judicial Risk Assessment
As bail reform debates continue, jurisdictions have adopted, revised or dropped pretrial tools. Evidence on whether the tools reduce detention without raising crime is mixed, and it depends heavily on how judges and policies use the scores. Likely directions include more public validation studies, transparent or open-source models, and limits on using algorithmic scores in sentencing. On generative AI, courts are issuing guidance for judges and clerks who use chatbots to draft and summarize. That guidance generally stresses confidentiality and keeps humans responsible for decisions.
現実世界の実装
A pretrial services officer runs the Public Safety Assessment, which scores failure to appear and new criminal activity from factors based on the defendant's age and criminal history. The officer gives the report to the judge before the first bail hearing.
In State v. Loomis (2016), the Wisconsin Supreme Court allowed a sentencing judge to consider a COMPAS score. It required written warnings about the tool's limits and said the score could not be the deciding factor.
A county court uses software to sort incoming e-filings by case type and flag incomplete ones. It also sends automated text reminders about hearing dates, which studies have linked to fewer missed appearances.
A state's pretrial agency revalidates its risk tool on local data and finds that its 'high risk' cutoff flags more people than intended. The agency adjusts the threshold through a public policy review.
リスクとガードレール
規制要件により、強力なプロトタイプが無効になる可能性があります。
過去のデータには、特定のコミュニティに害を及ぼすバイアスがコード化されている可能性があります。
レガシー システムでは、統合のボトルネックや隠れたコストが発生する可能性があります。
実装ロードマップ
問題の枠組みから評価まで、各分野の専門家を巻き込みます。
起動前に監査証跡とドキュメントを設計します。
コンプライアンスと安全義務を早期に検証します。
明確な停止基準とロールバック基準を使用して、段階的にロールアウトします。
探検を続けましょう
Free newsletter
Get the daily AI briefing
Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.
One email each weekday. Unsubscribe in one click. We never sell or share your address.
Test yourself
Take the AI in Courts and Judicial Risk Assessment quiz
Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.
Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation
よくある質問
What is AI in Courts and Judicial Risk Assessment?
Judicial risk assessment tools are statistical models that estimate how likely a defendant is to miss court or be rearrested. Judges use the scores in bail, sentencing and parole decisions. The tools can make decisions more consistent and reduce unnecessary detention, but they raise hard questions about fairness, transparency and due process, as the debate over the COMPAS tool showed.
What does the Public Safety Assessment (PSA) produce?
The PSA uses factors based on age and criminal history, with no interview, to produce separate pretrial scores. It estimates risk before trial. It says nothing about guilt.
What did ProPublica's 2016 analysis of COMPAS in Broward County report?
ProPublica focused on false positives: people labeled higher risk who did not go on to reoffend. That error rate was much higher for Black defendants.
How did Northpointe respond to ProPublica's findings?
Northpointe defended the tool as calibrated. The same score meant roughly the same risk for each group. This is a different fairness definition from equal error rates.
Why can a risk score generally not be calibrated and have equal false positive and false negative rates across groups at the same time?
Research by Kleinberg, Chouldechova and others showed that when base rates differ, these fairness criteria conflict except in trivial cases. So choosing which one to prioritize is a policy decision.
What did the Wisconsin Supreme Court decide in State v. Loomis (2016)?
The court allowed limited use with cautions, balancing a potentially useful tool against due process concerns about a proprietary model.
学び続ける
関連ガイド
このトピックのために選ばれたその他のガイド