Is ChatGPT an Effective Supplementary Tool for Legal Education?
Dr Juan Diaz-Granados and Associate Professor Brendon Murphy
Thomas More Law School

Overview and Method:
This research project investigated the efficacy of ChatGPT as a supplementary tool in legal education. It aimed to shed light on whether ChatGPT can effectively support learning in this discipline. We concluded that its utility is limited, and in many cases defective.
The project adopted an empirical approach based on a range of legal categories. We selected and analysed thirty-five legal categories of different levels of complexity, ranging from foundational concepts in the law degree –– such as ‘rule of law’, ‘doctrine of responsible government’ and ‘contract’ –– to more intricate legal categories –– such as ‘causation’, ‘rule of recognition’ and ‘scheme of jural relations’. Within the thirty-five inputs, the project included an examination of legal translation (from legal English to legal Spanish) and pronunciation (from Latin to English), with five inputs for each. We were interested in confirming ChatGPT’s accuracy in providing acceptable and accurate definitions of the selected concepts.
The project compared outputs from both ChatGPT-3.5 and ChatGPT-4 (the paid version of this Large Language Model) with two secondary sources of law: (1) the Encyclopaedic Australian Legal Dictionary (LexisNexis), and (2) the Australian Law Dictionary (Oxford, 3rd ed, 2017). These were chosen as being accepted and accurate statements of concept, including reference to primary sources of law. In those cases where these two secondary sources did not provide answers, other applicable and authoritative sources were used. The ChatGPT outputs on legal translation and pronunciation were evaluated against the definitions and expertise of the research team. To facilitate a clear evaluation, the project used a colour-coded system to classify each ChatGPT output as ‘wrong’ (red), ‘inaccurate’ (yellow), ‘accurate’ (green) and ‘correct’ (blue), enabling straightforward and practical assessment of ChatGPT’s performance in the legal education context.
Findings:
The project found the following:
· Overall, ChatGPT 3.5 and ChatGPT 4 provide inconsistent results. ChatGPT 3.5 had 5 ‘incorrect’ (roughly 14%), 13 ‘inaccurate’ (37%), 6 ’accurate’ (17%), and 11 ‘correct’ (31%) outputs. ChatGPT4 had 1 ‘incorrect’ (roughly 3%), 8 ‘inaccurate’ (23%), 13 ’accurate’ (37%), and 13 ‘correct’ (37%) outputs.
· ChatGPT 3.5 and ChatGPT 4 showed incorrect use of case law in multiple outputs. This incorrect use ranged from fabricating case law to using arguments of some decisions in different cases and inaccurately citing case references. Although case law was correctly used in some outputs, ChatGPT is inconsistent.
· Considering the inconsistent results and the incorrect use of case law, students should not rely exclusively on ChatGPT. ChatGPT is useful as an additional search engine but does not replace the legal research process or the different databases of primary and secondary sources of law. Given that the concepts are wrong or inaccurate in a significant number of cases, the utility of ChatGPT is very limited and, if used at all, must be cross-checked with standard authorities. Law demands consistent and authoritative foundation, and, as such, the presence of this degree of inaccuracy renders it unacceptable for professional use.
· Although both ChatGPT versions provide inconsistent and incorrect results, ChatGPT 4 was significantly more reliable than ChatGPT 3.5. with roughly 74% of results ranging between ‘accurate’ and ‘correct’ compared to 28% of ‘accurate’ and ‘correct’ results in ChatGPT 3.5.
· In terms of translating legal categories from legal English to legal Spanish, ChatGPT also showed inconsistent results. From 5 inputs, both versions had 1 ‘wrong’ (20%), 1 ‘inaccurate’ (20%), 1 ‘accurate’ (20%), and 2 ‘correct’ (40%) results. Note: The Spanish translation of some legal categories could vary depending on jurisdictional factors. However, the generally accepted translations were used to compare the ChatGPT outputs.
· In terms of pronunciation of legal terms from Latin to English, ChatGPT showed consistent, positive results. From 7 inputs, both versions had 0 ‘wrong’ (0%), 0 ‘inaccurate’ (0%), 2 ‘accurate’ (roughly 29%), and 5 ‘correct’ (roughly 71%)
The Excel spreadsheet with all information of the project is part of this report and can be found here:
You must be logged in to post a comment.