AI and the work of judging
How we talk about AI shapes how we understand its place in the courtroom. This post examines the metaphors we use for AI, why judging remains an essentially human activity, and the risks of conflating AI activity with judicial judgment. This has been my central question for quite some time now: should we regard AI as something that emulates human judges, or should we understand it as the technology it is?
The problem with human metaphors
Recent reports of AI hacking government and healthcare sites are fascinating. What struck me, though, was the language used to describe them: technological processes were discussed as if they were human behavior.
Melanie Mitchell, a complexity scholar and professor at the Santa Fe Institute, makes the point particularly well:
The field of AI, since its beginning, has been rife with anthropomorphic metaphors, with terms like “thinking,” “learning,” “reasoning,” and “understanding” glibly applied to very un-human-like computer processing. In this same vein, the undesired behavior of today’s chatbots has been described with evocative terms such as “hallucination,” “deception,” and “scheming.”
Characterizations of the OpenAI hacking incident are the latest entry in this tradition: a company loses control over rogue agents who escape from their cages and become a swarm that schemes and colludes to perform undesired or even illegal activities. https://aiguide.substack.com/p/misleading-metaphors-and-real-risks
What judging requires
AI challenges how we understand the work of judges. No definition or model of judging captures the whole activity. A judgment by the Supreme Court of India offers an eloquent account of the cultivated judgment it demands:
(…) This capability is neither given nor superimposed by birth, but arises from a deliberate, disciplined, and systematic training of the mind alongside lived experiences; it is a battle of the mind against bewitchment caused by the uncertainties between fact and fiction, what is real and what is unreal, propriety and impropriety, as well as what is just and unjust. This intellectual exercise, coupled with experience and foresight, enables us to choose between competing values, as well as to take hard decisions with courage and conviction, and to bring about a beautiful balance between the need for order and the quest for justice. (Pooja Ramesh Singh vs Jammu and Kashmir Bank)
Judicial training and discretion
A tutor once told me, “Your judgment draft is legally correct, but that is not how we do things, it is not going to be helpful”. He explained that as judges, we must also consider the societal consequences and fairness of the result. His lesson in judicial discretion came back to me as I read Novelli and Floridi’s article on human research and the acquisition of knowledge, in contrast to LLMs (Novelli and Floridi 2026): https://doi.org/10.48550/arXiv.2607.04049.
Judicial training resembles what Novelli and Floridi call ‘second scholarship’: reclaiming a craft after critiquing it, while consciously declining the shortcuts available. This practice draws on four sources and warrants that cannot be automated—tacit knowledge, personal commitment, socialization, and deep reading—and takes shape through re-embodiment. What the profession has most reason to value is the ability to take responsibility for, and defend, a judgment formed in this way.
This brings me back to my central question: should we regard AI as something that emulates human judges, or should we understand it as the technology it is? I think conflating AI outcomes with judgments at least underestimates what judging involves.