The Challenge of Legal Reasoning and IRAC Automation

Legal reasoning presents unique challenges for artificial intelligence systems, requiring a combination of factual understanding, rule identification, and logical application. The IRAC (Issue, Rule, Application, Conclusion) framework is the problem-solving approach widely used by legal professionals to determine underlying legal issues, extract and transform facts, and conduct legal reasoning that leads to a legal conclusion. However, recent studies show that large language models struggle with this task. On average, ChatGPT draws wrong conclusions on approximately 50% of legal scenarios, and even when conclusions are correct, there are mistakes in the intermediate reasoning steps. These models also fail to cite correct legal rules when analyzing the majority of legal scenarios, highlighting the gap between AI performance and the standards required for professional legal practice.

To address these limitations, researchers are developing specialized benchmarks and knowledge bases that support structured legal reasoning. LegalSemi, for example, comprises 54 legal scenarios rigorously annotated by legal experts based on the IRAC framework from Malaysian Contract Law, accompanied by a semi-structured knowledge base (SKE). The SKE represents legal concepts, court cases, legal rules, and interpretations in a structured format that enables efficient retrieval and reasoning. Experimental results demonstrate that incorporating the SKE improves issue identification by over 21.4% and rule retrieval significantly across all evaluated LLMs. This approach mitigates the linguistic disparity between legalese and everyday language, helping LLMs comprehend underlying legal knowledge more accurately.

Several benchmarks have been developed to advance legal reasoning capabilities, including SARA for statutory reasoning, CaseHOLD for identifying holding statements in US case law, and LegalBench with over 100 tasks for assessing legal reasoning across multiple dimensions. These benchmarks assess capabilities including statutory interpretation, contract analysis, and case outcome prediction. However, a critical gap remains: no existing dataset explicitly annotates complete reasoning steps or adopts a comprehensive IRAC framework. This limitation is significant, as reasoning capabilities are fundamental to solving legal problems effectively. Current research aims to address this by developing frameworks that explicitly model the legal reasoning process through structured knowledge bases and explicit reasoning steps, paving the way for AI systems that can participate meaningfully in legal analysis.

The Rise of Legal General Intelligence and the Human Gap

The concept of legal intelligence is evolving far beyond simple legal research automation. Researchers are now defining “legal general intelligence” (Legal GI) as the capacity of artificial intelligence to perform with expert-level ability across complex legal contexts, including the interpretation of provisions, sound inference, conflict resolution between multiple legal domains, and making normatively binding judgments in ethically sensitive contexts. This represents a fundamental shift from task-oriented tools to systems that can participate in the normative structure of legal systems themselves. Benchmarks like LexGenius have emerged to systematically evaluate this capability, moving beyond outcome-focused assessments to examine the underlying reasoning processes.

The gap between AI and human legal professionals remains substantial. Recent evaluations of 12 state-of-the-art large language models on legal intelligence benchmarks reveal significant disparities across legal abilities, with even the best-performing models lagging behind human legal experts. The challenge is particularly acute in what researchers call “soft legal intelligence”—areas like ethical judgment, law-morality boundaries, and societal impact assessment. Traditional legal benchmarks have focused on technical tasks while overlooking these critical dimensions, creating a false impression of AI competence.

To address these limitations, frameworks are being developed to structure legal intelligence evaluation across multiple dimensions, tasks, and abilities. These frameworks draw on Bloom’s Taxonomy of Educational Objectives, covering the cognitive hierarchy from remembering and understanding to creating, alongside modular models used in legal evaluations across countries. The goal is not just to test whether AI knows the law, but whether it can engage with the full complexity of legal practice. As these evaluation frameworks mature, they pave the way for legal AI to move “towards the second half of AI,” where systems demonstrate genuine understanding rather than pattern mimicry.