When the Hallucination Becomes the Law: How AI’s Greatest Flaw Is Forcing Justice to Define Itself
The mistake at the Income Tax Appellate Tribunal in Bengaluru was not dramatic. It did not involve a complex legal principle or a landmark constitutional question. The error was quiet, buried in the text of an order that cited judicial precedents that had never existed. The judges who wrote the order had not fabricated these precedents themselves. They had used an AI tool to assist with legal research, and the tool had done what generative AI systems do when they lack reliable information: it invented something plausible. The invented judgments were cited as if they were real. The order was issued. And then, in February 2025, someone checked the citations and discovered that the AI had produced hallucinations — confident, detailed, and entirely false legal references that looked authentic enough to pass through a court
This single mistake triggered a chain of causality that would reshape the Indian judiciary’s relationship with technology. The Bengaluru incident led directly to the Karnataka High Court calling for action against a civil court judge who had similarly relied on non-existent Supreme Court and Delhi High Court judgments. That case led to the Supreme Court taking notice of allegations that the National Company Law Tribunal and the National Company Law Appellate Tribunal had used fake judicial citations in insolvency proceedings involving Essel Infraprojects Ltd. And that chain of judicial embarrassment led, by June 2026, to the Supreme Court releasing draft regulations for the use of artificial intelligence in courts — perhaps the most comprehensive judicial AI governance framework attempted by any major jurisdiction in the world.
The chain reveals something important about how institutions learn. The mistake was not an anomaly. It was a symptom of a deeper problem with how AI systems work. Generative AI models do not store facts the way a database stores records. They store patterns of language. When asked to produce a legal citation, they do not retrieve a stored memory of a real case. They generate a sequence of words that matches the pattern of a legal citation. The result looks like a real judgment, reads like a real judgment, and sounds authoritative. But it is a statistical approximation of what a real judgment would look like, not a reference to an actual legal document. The AI was not lying. It was doing exactly what it was designed to do: produce text that matches the statistical distribution of its training data. The problem is that the training data for legal citations includes real cases, and the model learned to reproduce the pattern of those real cases without learning which cases were real and which were invented.
The Supreme Court’s proposed regulations address this problem directly. The draft framework, released under Chief Justice of India Surya Kant and chaired by Justice P.S. Narasimha, creates a legal architecture that treats AI hallucinations not as a technical glitch but as a systemic risk that requires structural safeguards. The regulations do not ban AI from legal practice. They do not treat the technology as inherently dangerous. Instead, they establish a principle of accountability that runs through every provision: the human who uses AI remains responsible for what the AI produces. A lawyer who files a pleading containing AI-generated false citations cannot escape liability by blaming the tool. The mistake belongs to the person who submitted it, not to the system that generated it.
This principle of human responsibility is the first link in a chain of reasoning that leads to the most consequential aspect of the regulations: the categorical prohibition on AI making judicial decisions. The draft regulations declare that AI must remain “strictly subservient” to human judgment and judicial authority. The power to determine questions of law, facts, and justice will continue to vest exclusively in judges. This is not merely a policy preference. It is a constitutional boundary drawn around the core function of the judiciary. The regulations prohibit AI from deciding cases, passing sentences, or determining judicial outcomes through algorithmic decision-making. Even where AI assists, its outputs must remain advisory and subject to independent judicial scrutiny. Judges cannot abdicate their responsibility by relying on machine-generated recommendations.
The prohibition extends beyond decision-making to prediction. The regulations bar AI from assessing flight risk, predicting recidivism, determining bail eligibility, evaluating the credibility of witnesses and parties, profiling litigants, or predicting future behavior. These restrictions reflect an understanding that algorithmic prediction in legal contexts carries risks that go beyond inaccuracy. Predictive algorithms trained on historical data encode the biases present in that data. If the criminal justice system has historically treated certain communities more harshly, an algorithm trained on that history will reproduce that harshness as a statistical pattern. The algorithm does not need to be programmed with racist intent. It simply learns from what exists. And what exists in many legal systems is a history of unequal treatment that the algorithm will faithfully replicate as predictive truth.
The chain of causality that began with the Bengaluru ITAT mistake now leads to a deeper question about what AI reveals about the nature of legal knowledge itself. The AI hallucinations that embarrassed the courts were not random errors. They were systematic failures of a particular kind of knowledge representation. Legal knowledge is not just a collection of facts. It is a network of relationships between principles, precedents, statutes, and interpretations. A legal citation is meaningful only because it connects a specific proposition to a specific authority that validates that proposition. When an AI generates a false citation, it creates a connection that looks real but has no authority behind it. The citation is a ghost — a pattern without substance.
This is where the Supreme Court’s regulations show their deepest understanding of the problem. The draft framework does not just prohibit false citations. It creates an entire governance architecture designed to ensure that AI systems used in courts are transparent, explainable, and auditable. The regulations propose an AI Content Verification Authority to develop standards and protocols for verifying generative AI outputs. They mandate technical and ethical impact assessments before any AI system is deployed, evaluating training data, risks of bias, hallucinations, cybersecurity vulnerabilities, and compliance with human oversight requirements. They require AI audits, AI registers, incident reporting systems, and annual transparency reports. High courts, tribunals, and commissions would be required to publicly disclose the AI systems they use, audit outcomes, and incidents recorded.
The architecture is elaborate because the problem is fundamental. AI systems that generate text cannot simply be told to stop hallucinating. Hallucination is not a bug in these systems. It is a feature of how they work. The same mechanisms that allow them to generate creative and novel outputs also allow them to generate false outputs. There is no way to have one without the other. The only solution is to create institutional processes that catch hallucinations before they cause harm, and to assign clear responsibility when they slip through.

The historical background to these regulations stretches back to 2019, when the Supreme Court constituted its first AI committee to explore how technology could be integrated responsibly into the justice system. That committee was reconstituted in December 2025 under Chief Justice Surya Kant, with a mandate to steer the adoption, development, and deployment of AI tools across the higher judiciary and subordinate courts. The committee’s work accelerated after the series of hallucination incidents in early 2025 and 2026, which demonstrated that the problem was not hypothetical. Real cases were being decided based on fake legal authorities. Real litigants were being affected by errors that no human had caught because no human had checked.
The regulations also address a deeper historical pattern: the tendency of institutions to adopt technology first and regulate later. The internet was adopted by courts before clear rules existed for electronic filing and digital evidence. Social media transformed legal practice before courts understood how to handle online harassment of judges and jurors. AI adoption was following the same pattern until the Supreme Court intervened. The draft regulations represent an attempt to reverse the usual sequence: establish the governance framework before large-scale deployment makes regulation difficult or impossible.
The chain of causality now extends beyond India. The Supreme Court’s draft regulations are being studied by judicial systems around the world that face the same problem. Courts in the United States, the United Kingdom, and Australia have all encountered lawyers citing AI-generated false cases. The responses have been ad hoc: sanctions in individual cases, warnings from judges, and bar association guidance. No major jurisdiction has attempted the comprehensive approach that India is proposing. The Indian framework is significant not just for what it does domestically, but for what it offers as a model for other countries grappling with the same issues.
The regulations permit AI to assist with legal research, citation verification, summarization of pleadings and judgments, translation, transcription, drafting assistance, scheduling, record management, and case administration. AI-powered chatbots may help litigants understand procedures and access court services. Accessibility tools for persons with disabilities are specifically encouraged. The framework creates a “presumption in favor of responsible AI adoption,” signaling that courts should actively explore technologies capable of reducing delays and improving judicial administration.
But the permitted uses are carefully bounded. AI-generated material cannot be treated as independent evidence unless its AI-generated nature is fully disclosed. Opaque or unexplainable AI systems are barred from matters affecting legal rights or personal liberty. Private vendors cannot participate in court AI systems without prior approval, and contracts must contain provisions on ownership of court data, restrictions on data usage, audit rights, transparency obligations, and liability for harm.
The most forward-looking aspect of the draft is the institutional structure proposed for governing AI in courts. The regulations envisage a permanent apex body at the national level to establish standards, approve AI systems, and supervise AI adoption across the judiciary. Supporting this body would be specialized committees dealing with judicial applications, technology, infrastructure, finance, cybersecurity, and data management. A dedicated Centre of Research and Excellence on Artificial Intelligence (CoRE-AI) is proposed to provide technical and legal support. Every high court would have its own AI Committee and AI Secretariat, headed by judicial officers and assisted by technology experts.
This institutional architecture reflects an understanding that AI governance cannot be a one-time policy decision. It must be an ongoing process of oversight, adaptation, and learning. The technology will change. New capabilities will emerge. New risks will appear. The institutions created by the regulations are designed to evolve with the technology, not to freeze a particular approach in time.
The chain of causality that began with a single hallucinated citation in Bengaluru has now produced a framework that addresses not just hallucinations but the entire relationship between human judgment and machine assistance. The regulations recognize that AI can help judges work faster, but they insist that it can never become a judge. They acknowledge that AI can assist advocates, but they insist that responsibility remains human. They permit innovation, but they require transparency.
And this brings us to the question a child might ask about all of this. The child has heard about the AI that invented fake court cases, and the child wants to understand why this matters. The child asks: If the AI made up a case that sounded real, and the judge believed it, and the judge made a decision based on that made-up case, then who really decided the case — the judge or the AI?
The question is deeper than it appears. It goes to the heart of what it means to decide. A judge who relies on a false citation produced by an AI is not exercising judgment in any meaningful sense. The judge is ratifying a decision that the AI has effectively made by providing the legal authority for it. The AI did not decide the outcome directly, but it shaped the reasoning that led to the outcome. The judge became a conduit for the AI’s output rather than an independent decision-maker.

The Supreme Court’s regulations answer this question by drawing a bright line: AI may assist, but it cannot decide. The line is simple in principle but difficult to maintain in practice. Every time a lawyer uses AI to draft a pleading, every time a judge uses AI to research a precedent, every time a court uses AI to summarize a complex argument, the line between assistance and decision-making becomes harder to see. The regulations do not pretend that the line is easy to draw. They create institutions and processes to keep drawing it, case by case, as the technology evolves.
The child’s question also reveals something about the nature of legal authority. A legal system rests on the assumption that decisions are made by humans who can be held accountable for their reasoning. A judge who makes a mistake can be appealed. A lawyer who misleads the court can be sanctioned. But an AI that makes a mistake cannot be held accountable. It cannot be sanctioned. It cannot explain its reasoning. It can only produce more text. The legal system’s response to AI hallucinations is not just about preventing errors. It is about preserving the structure of accountability that makes the legal system legitimate.
The chain of causality that began with a mistake in Bengaluru has now reached its logical conclusion: a comprehensive framework for governing AI in courts that treats the technology as a tool, not a substitute for human judgment. The regulations do not solve the problem of AI hallucinations. They create the conditions under which the problem can be managed. They establish the principle that humans remain responsible for what machines produce. They build the institutions that will oversee this responsibility as the technology changes.
The child’s question — who really decided the case, the judge or the AI — does not have a simple answer. But the Supreme Court’s regulations provide the framework within which the question can be asked, debated, and answered in each specific case. That is perhaps the most important thing the regulations do. They ensure that the question of who decides remains a question that humans must answer, not a question that machines can decide by default. The framework does not eliminate the tension between human judgment and machine assistance; it forces the legal system to confront it, case by case, as the technology evolves.
Sources
1. Income Tax Appellate Tribunal, Bengaluru
