Turing Test in Italian Legal Exams: LLMs Excel in Argumentation but Fail in Notarial Planning
Researchers tested top AI models on difficult Italian professional exams for lawyers, judges, and notaries. While some AIs were as good as or better than top humans at arguing legal points, they all failed the notary exam because they couldn't handle complex, goal-oriented planning.
The Gap Between Recall and Application
Can a machine truly "think" like a legal professional? While Large Language Models (LLMs) have shown remarkable ability to solve STEM problems, the legal domain introduces a different kind of friction. In law, accuracy isn't just about retrieving the right statute. It is about interpreting nuances and navigating conflicting precedents. It also requires applying abstract principles to messy, real-world scenarios.
The central tension in current legal AI research is the gap between "rule recall" (knowing what the law says) and "rule application" (using that law to solve a specific problem). Many existing benchmarks focus on multiple-choice questions or short snippets of text. These essentially test pattern matching. This study seeks to push past that. It tests LLMs on long-form, complex legal reasoning through a blind "Turing Test" involving high-stakes professional examinations.
The Italian Professional Framework
To stress-test these models, the authors utilized three distinct tiers of the Italian civil law system. Each tier represents a different cognitive profile:
- The Bar Exam: This focuses on adversarial reasoning (arguing a side to win a dispute). Candidates must draft legal summons (formal documents starting a lawsuit). They must use persuasive argumentation to defend a client's interests against an opponent.
- The Judicial Exam: This is a doctrinal task (analyzing legal theory and principles). It requires an impartial, systemic analysis of legal issues. The candidate must navigate competing academic theories and judicial interpretations without taking sides.
- The Notary Exam: This is the most application-oriented. A notary must draft legal acts—such as wills or property deeds. These must satisfy strict formal requirements. They must also achieve the specific, long-term goals of the parties involved.
Think of it like the difference between a debater (Bar), a philosopher (Judicial), and an architect (Notary). The debater wins by arguing a position. The philosopher explains the underlying principles. The architect must ensure the entire structure is both functional and meets all regulatory standards.
Testing the Frontiers
The researchers conducted a blind experiment using four state-of-the-art, "out-of-the-box" models. Note that the model names used in this paper (Claude 4 Opus, GPT-5, DeepSeek R1, and Gemini 2.5 Pro) reflect the specific experimental context of the study.
To ensure a rigorous "Turing Test," the authors had expert examiners evaluate the AI-generated papers alongside the highest-scoring human essays from actual exams. The examiners were not told which papers were human and which were machine-generated.
The study reveals a stark divergence in performance based on the nature of the task. In the Bar and Judicial exams, the models showed impressive competence. The authors report that Gemini 2.5 Pro actually exceeded human performance in several categories. Specifically, in the Bar exam, Gemini achieved a score of 79. This surpassed the top human candidate's score of 62. In the Judicial exam, Gemini earned 21/24. This outperformed the human benchmark of 18/24. The paper notes that these models demonstrated advanced adversarial reasoning. They also showed a sophisticated command of the "doctrinal landscape" (the complex web of academic and judicial opinions).
However, the results shifted dramatically during the Notary exam. The authors find a "performance ceiling" here. While the models could argue or explain, they could not plan. In the inter vivos (living) assignment, the human candidate achieved a total score of 42. The best-performing model, Gemini 2.5 Pro, managed only 10. Even in the mortis causa (death-related) assignment, where GPT-5 performed relatively better, all models fell significantly below the human standard.
The failure in the notary domain is not due to a lack of language skill. Instead, it stems from a lack of "goal-directed legal planning." The authors categorize these failures into a specific taxonomy: legal-source failures (incorrect rules), reasoning failures (internal inconsistency), pertinence failures (failing to address the core issue), lexical failures (improper terminology), and formal failures (structural defects).
Identifying the Planning Deficit
What these results tell us is that "legal competence" is not a monolithic trait. High performance in argumentative or doctrinal tasks does not automatically translate to the ability to design complex, risk-sensitive legal instruments.
The study suggests that LLMs currently excel at synthesizing existing knowledge. They effectively organize what is already "known" in their training data. However, they struggle when they must coordinate multiple moving parts to reach a specific objective.
The models also exhibited a notable "gullibility" when faced with legal traps. In the inter vivos assignment, the scenario involved a property transfer to a person who was deaf-mute but able to communicate via reading and writing. Under Italian notarial law, a sign-language interpreter is mandatory to prevent nullity. All models incorrectly stated that an interpreter was not needed. This illustrates how models can be misled by specific situational nuances despite knowing the general legal sources.
Limits of the Experiment
The authors are careful to note several boundaries to their findings. First, the LLMs were granted access to the internet during the test. Human candidates were traditionally restricted to the Italian civil code. This creates an asymmetry in how information is retrieved. Second, the models were tested "out-of-the-box." This means they had not undergone specialized legal fine-tuning or used Retrieval-Augmented Generation (RAG; a technique where a model consults specific, external documents before answering).
Finally, the study is localized to the Italian civil law system. Because legal reasoning is deeply tied to specific jurisdictions and linguistic nuances, the results might not generalize to common law systems or other legal traditions.
The takeaway for the legal industry is one of cautious utility. The paper suggests that while LLMs may soon serve as powerful assistants for drafting persuasive arguments or summarizing legal doctrine, they are not yet ready for autonomous roles. They are not yet ready for high-stakes planning or formal instrument drafting. In the hands of a professional, they are powerful research tools. In isolation, they remain prone to "gullibility" and planning collapses.
Figures from the paper
How this was made
Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: academic_accessible
Template: explainer
Refinement: 1
Pipeline: forge-1.1
Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 85% (passed)
Claims verified: 15 / 15
Model: nvidia/Gemma-4-26B-A4B-NVFP4
NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 112,001
Wall-time: 315.3s
Tokens/s: 355.2