Was ist passiert?
Emanuele Ratti and Lena Zuchowski posted a paper, accepted for publication in Studies in the History and Philosophy of Science, that treats the often-invoked comparison between clinical translation and machine-learning development as a 'generative analogy' and builds from it a reliabilist account of when ML systems are epistemically warranted.
A paper titled 'What Can Artificial Intelligence Learn from Medicine? Generative Analogies and Reliable Machine Learning Systems,' by Emanuele Ratti and Lena Zuchowski, was submitted to arXiv on 18 August 2026 and is listed as arXiv:2608.18186. The listing states the paper has been accepted for publication in Studies in the History and Philosophy of Science, and asks that citations refer to the published version. It is filed under Machine Learning (cs.LG) with cross-listings to Artificial Intelligence (cs.AI) and Computers and Society (cs.CY). The version posted is v1, a 416 KB PDF.
The paper's starting point, as described in its abstract, is a problem of justification rather than performance. Machine learning has been widely and, the authors say, to an extent successfully implemented in medicine, but uncertainties surrounding it have made it hard to establish the bases of its epistemic and methodological warrants — that is, what entitles anyone to treat a system's outputs as trustworthy. Existing literature, the authors note, has drawn a parallel between medicine and machine learning, suggesting that standards for ML should be modeled on the standards used in clinical translation, the staged process by which a candidate intervention moves from laboratory finding to approved clinical use.
The paper's contribution is to make that parallel precise instead of gestural. Drawing on tools from Hesse's work on analogical reasoning in science, the authors characterise the relationship between clinical translation and the building of ML systems as a generative analogy — a comparison that does argumentative work and yields new claims, rather than merely illustrating a point. They then identify the specific epistemic and methodological warrants of clinical translation that, in their account, are usually only alluded to when the analogy is invoked, and argue for the sense in which those warrants carry over analogically to machine learning.
The final move is interpretive: the authors read the warrants of clinical translation in reliabilist terms, and use that reading to propose what they call a new form of ML reliabilism. They present it as distinct from existing reliabilist accounts in the philosophy of AI, though compatible with them. Reliabilism, in epistemology, locates justification in the reliability of the process that generated a belief rather than in the believer's grasp of supporting reasons.
Several things are not established by the material reviewed here, which is the arXiv abstract and listing page rather than the full text. The abstract does not enumerate which warrants of clinical translation the authors select, which regulatory or trial frameworks they treat as the reference case, or which competing reliabilist accounts they distinguish themselves from. No experiments, benchmarks, datasets, clinical results or case studies are described. The journal issue, page numbers and publication date are not stated, and the arXiv page does not list institutional affiliations. Acceptance at a journal is reported by the listing itself; the posted preprint is not a peer-reviewed artifact, and the published version may differ.
Lesen Sie die Primärquelle: arxiv.org ↗
Warum es wichtig ist
The 'clinical trials for AI' comparison already circulates in policy and assurance debates, usually without stating what it actually implies. Making the analogy explicit locates the warrant for trusting a model in the reliability of the process that produced it rather than in explanations of the model itself — a different target for auditors and regulators than interpretability.
The comparison this paper formalises is already in wide circulation. Calls for 'clinical trials for AI,' staged deployment, phased evaluation and post-market monitoring appear regularly in policy discussion, procurement language and assurance proposals, usually as an appeal to intuition rather than a worked-out argument. Analogies imported loosely tend to carry unexamined assumptions with them: which parts of the medical apparatus are being borrowed, and which are quietly left behind, often goes unsaid. A paper that states the mapping explicitly makes it possible to argue about, and to reject in specific places.
The reliabilist framing has a practical edge. If the warrant for trusting a model's output comes from the reliability of the process that produced and validated it, then the evidence that matters is procedural: how data were collected and partitioned, how the system was tested against populations it will actually meet, what monitoring continues after deployment. That is a different object of scrutiny than the one much of the current debate focuses on, where trust is sought through explanation and interpretability of the model's internals. The two are not mutually exclusive, but they direct auditing effort at different places and generate different documentation.
For people who buy, deploy or oversee medical machine learning, the distinction bears on what a vendor should be asked to show. A process-reliability standard points toward staged evidence, prospective validation and surveillance obligations; an explanation-centred standard points toward model transparency. The paper, on the evidence of its abstract, argues at the level of what would justify trust rather than supplying an instrument. It does not appear to offer thresholds, a checklist, an evaluation protocol or a certification scheme, and readers looking for one will not find it here.
The analogy also has limits the paper's own framing implies. Clinical translation's warrants are not free-floading epistemic virtues; they are sustained by institutions — ethics review, trial registration, regulatory gatekeeping, adverse-event reporting, professional liability — that have no complete counterpart for machine-learning systems, and that took decades to build. Medicine's own standards are contested from inside, with continuing disputes about replication, external validity and how well trial populations represent patients. A generative analogy is generative in part because of where it breaks; how much weight the borrowed standards can carry is exactly what the full text would need to settle, and cannot be judged from an abstract.
There is also a scope question worth flagging. The paper is framed around medicine, which is unusually well supplied with translational machinery. Whether the same warrants transfer to machine learning in domains with thinner institutional scaffolding — hiring, credit, public administration — is not something the abstract claims, and should not be assumed on its behalf.
Was Sie als nächstes sehen sollten
Whether the published journal version changes the argument, whether other philosophers of AI accept 'ML reliabilism' as distinct from existing reliabilist accounts, and whether anyone converts the framework into concrete evaluation or documentation criteria that developers and regulators could apply.
The first thing to track is the published version. The listing directs readers to cite the journal article rather than the preprint, so the definitive text — and any changes made in review to the characterisation of the warrants or the reliabilist claim — will appear in Studies in the History and Philosophy of Science. The issue, date and final DOI are not yet stated on the arXiv page.
The second is reception among philosophers of AI. The paper positions ML reliabilism as distinct from, though compatible with, existing reliabilist accounts. That claim of distinctness is the kind of thing specialists contest directly, and the useful signal will be whether subsequent work adopts the label and the mapping, argues the distinction collapses into an existing account, or attacks the analogy at a specific joint — for example by arguing that machine-learning pipelines lack a stable analogue of the trial phases that give clinical warrants their force.
The third is whether anything operational follows. Conceptual accounts of assurance become consequential when they are converted into criteria: documentation requirements, staged-release conditions, monitoring duties, or evidence standards used by standards bodies, health systems and regulators. Watch for follow-up work applying the framework to actual deployed medical ML systems, and for whether assurance practitioners cite it when justifying process-based rather than explanation-based evidence.
Finally, watch the counter-argument. If a process-reliability account is taken to displace interpretability as the ground of trust, expect pushback from researchers who hold that clinicians and patients are owed reasons, not only reliable procedures — and from those who argue that the institutional preconditions of clinical warrants simply do not exist for software that is updated continuously. Evidence that the analogy misleads in a concrete case would matter as much to this debate as evidence that it holds.


