뉴스로 돌아가기
혁신AI Understanding 브리핑

Philosophy paper argues machine learning should borrow its standards of proof from clinical translation

A preprint accepted by Studies in the History and Philosophy of Science treats the comparison between medicine and machine learning as a formal analogy rather than a slogan, and uses it to sketch a process-based account of when an ML system deserves trust. It is conceptual: no experiments, thresholds or checklists.

7 min readRead the primary source
Source-page capture accompanying Philosophy paper argues machine learning should borrow its standards of proof from clinical translation
기본 소스 문서녹음된 소스
출판사
arxiv.org
소스 링크
arxiv.orghttps://arxiv.org/abs/2608.18186
소스 유형
기본 문서 — 우리가 직접 읽는 공식 발표, 논문, 서류 또는 자사 페이지입니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

기계 학습(ML)
시스템이 데이터로부터 패턴을 학습하고 시간이 지남에 따라 개선될 수 있도록 하는 방법입니다.
인공지능(AI)
패턴 인식, 추론, 언어 또는 의사 결정이 필요한 작업을 수행하는 시스템 구축의 광범위한 분야입니다.
인용
모델의 주장을 뒷받침하기 위해 모델의 응답에 포함된 소스 구절이나 문서에 대한 참조입니다.
자신을 테스트해 보세요AI 윤리 퀴즈

무슨 일이 일어났나요?

Emanuele Ratti and Lena Zuchowski posted a paper, accepted for publication in Studies in the History and Philosophy of Science, that treats the often-invoked comparison between clinical translation and machine-learning development as a 'generative analogy' and builds from it a reliabilist account of when ML systems are epistemically warranted.

A paper titled 'What Can Artificial Intelligence Learn from Medicine? Generative Analogies and Reliable Machine Learning Systems,' by Emanuele Ratti and Lena Zuchowski, was submitted to arXiv on 18 August 2026 and is listed as arXiv:2608.18186. The listing states the paper has been accepted for publication in Studies in the History and Philosophy of Science, and asks that refer to the published version. It is filed under Machine Learning (cs.LG) with cross-listings to Artificial Intelligence (cs.AI) and Computers and Society (cs.CY). The version posted is v1, a 416 KB PDF.

The paper's starting point, as described in its abstract, is a problem of justification rather than performance. Machine learning has been widely and, the authors say, to an extent successfully implemented in medicine, but uncertainties surrounding it have made it hard to establish the bases of its epistemic and methodological warrants — that is, what entitles anyone to treat a system's outputs as trustworthy. Existing literature, the authors note, has drawn a parallel between medicine and machine learning, suggesting that standards for ML should be modeled on the standards used in clinical translation, the staged process by which a candidate intervention moves from laboratory finding to approved clinical use.

The paper's contribution is to make that parallel precise instead of gestural. Drawing on tools from Hesse's work on analogical reasoning in science, the authors characterise the relationship between clinical translation and the building of ML systems as a generative analogy — a comparison that does argumentative work and yields new claims, rather than merely illustrating a point. They then identify the specific epistemic and methodological warrants of clinical translation that, in their account, are usually only alluded to when the analogy is invoked, and argue for the sense in which those warrants carry over analogically to machine learning.

The final move is interpretive: the authors read the warrants of clinical translation in reliabilist terms, and use that reading to propose what they call a new form of ML reliabilism. They present it as distinct from existing reliabilist accounts in the philosophy of AI, though compatible with them. Reliabilism, in epistemology, locates justification in the reliability of the process that generated a belief rather than in the believer's grasp of supporting reasons.

Several things are not established by the material reviewed here, which is the arXiv abstract and listing page rather than the full text. The abstract does not enumerate which warrants of clinical translation the authors select, which regulatory or trial frameworks they treat as the reference case, or which competing reliabilist accounts they distinguish themselves from. No experiments, benchmarks, datasets, clinical results or case studies are described. The journal issue, page numbers and publication date are not stated, and the arXiv page does not list institutional affiliations. Acceptance at a journal is reported by the listing itself; the posted preprint is not a peer-reviewed artifact, and the published version may differ.

소스 세부정보: arxiv.org

왜 중요한가요?

The 'clinical trials for AI' comparison already circulates in policy and assurance debates, usually without stating what it actually implies. Making the analogy explicit locates the warrant for trusting a model in the reliability of the process that produced it rather than in explanations of the model itself — a different target for auditors and regulators than interpretability.

The comparison this paper formalises is already in wide circulation. Calls for 'clinical trials for AI,' staged deployment, phased evaluation and post-market monitoring appear regularly in policy discussion, procurement language and assurance proposals, usually as an appeal to intuition rather than a worked-out argument. Analogies imported loosely tend to carry unexamined assumptions with them: which parts of the medical apparatus are being borrowed, and which are quietly left behind, often goes unsaid. A paper that states the mapping explicitly makes it possible to argue about, and to reject in specific places.

The reliabilist framing has a practical edge. If the warrant for trusting a model's output comes from the reliability of the process that produced and validated it, then the evidence that matters is procedural: how data were collected and partitioned, how the system was tested against populations it will actually meet, what monitoring continues after deployment. That is a different object of scrutiny than the one much of the current debate focuses on, where trust is sought through explanation and interpretability of the model's internals. The two are not mutually exclusive, but they direct auditing effort at different places and generate different documentation.

For people who buy, deploy or oversee medical machine learning, the distinction bears on what a vendor should be asked to show. A process-reliability standard points toward staged evidence, prospective validation and surveillance obligations; an explanation-centred standard points toward model transparency. The paper, on the evidence of its abstract, argues at the level of what would justify trust rather than supplying an instrument. It does not appear to offer thresholds, a checklist, an evaluation protocol or a certification scheme, and readers looking for one will not find it here.

The analogy also has limits the paper's own framing implies. Clinical translation's warrants are not free-floading epistemic virtues; they are sustained by institutions — ethics review, trial registration, regulatory gatekeeping, adverse-event reporting, professional liability — that have no complete counterpart for machine-learning systems, and that took decades to build. Medicine's own standards are contested from inside, with continuing disputes about replication, external validity and how well trial populations represent patients. A generative analogy is generative in part because of where it breaks; how much weight the borrowed standards can carry is exactly what the full text would need to settle, and cannot be judged from an abstract.

There is also a scope question worth flagging. The paper is framed around medicine, which is unusually well supplied with translational machinery. Whether the same warrants transfer to machine learning in domains with thinner institutional scaffolding — hiring, credit, public administration — is not something the abstract claims, and should not be assumed on its behalf.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
대화형 개념 확인+10 Points
AI Ethics Quiz

Impossibility results in algorithmic fairness (e.g. Kleinberg et al., Chouldechova) show what?

다음에 무엇을 볼 것인가

Whether the published journal version changes the argument, whether other philosophers of AI accept 'ML reliabilism' as distinct from existing reliabilist accounts, and whether anyone converts the framework into concrete evaluation or documentation criteria that developers and regulators could apply.

The first thing to track is the published version. The listing directs readers to cite the journal article rather than the preprint, so the definitive text — and any changes made in review to the characterisation of the warrants or the reliabilist claim — will appear in Studies in the History and Philosophy of Science. The issue, date and final DOI are not yet stated on the arXiv page.

The second is reception among philosophers of AI. The paper positions ML reliabilism as distinct from, though compatible with, existing reliabilist accounts. That claim of distinctness is the kind of thing specialists contest directly, and the useful signal will be whether subsequent work adopts the label and the mapping, argues the distinction collapses into an existing account, or attacks the analogy at a specific joint — for example by arguing that machine-learning pipelines lack a stable analogue of the trial phases that give clinical warrants their force.

The third is whether anything operational follows. Conceptual accounts of assurance become consequential when they are converted into criteria: documentation requirements, staged-release conditions, monitoring duties, or evidence standards used by standards bodies, health systems and regulators. Watch for follow-up work applying the framework to actual deployed medical ML systems, and for whether assurance practitioners cite it when justifying process-based rather than explanation-based evidence.

Finally, watch the counter-argument. If a process-reliability account is taken to displace interpretability as the ground of trust, expect pushback from researchers who hold that clinicians and patients are owed reasons, not only reliable procedures — and from those who argue that the institutional preconditions of clinical warrants simply do not exist for software that is updated continuously. Evidence that the analogy misleads in a concrete case would matter as much to this debate as evidence that it holds.

관련 가이드 및 퀴즈

AI 윤리AI 모델 설명AI란 무엇인가?알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.
이것이 유용하다고 생각하시나요?