Follow Us:

LLM in Tax and Compliance Practice: An Assessment of Auditability, Reliability and Professional Risk

Summary: The supplied content is a General News article examining the use of Large Language Models (LLMs) in tax and compliance practice. It states that LLMs can assist with research, drafting, judgment summaries, and client queries but identifies risks including hallucinations, outdated responses caused by training cutoffs, absence of a verifiable chain of custody, citation hallucinations, data privacy concerns, and the underrepresentation of Indian legal and regulatory material in training datasets. It notes that practitioners researching tax provisions, GST, judicial pronouncements, or regulatory developments may receive responses that do not reflect current law and emphasises that LLM outputs should be independently verified against primary sources. The article also discusses confidentiality concerns arising from the submission of client information to consumer LLM tools in light of the ICAI Code of Ethics and refers to documented incidents involving OpenAI, Samsung Electronics, the Italian data protection authority, and major financial institutions. It proposes a framework recommending that LLM outputs be treated as unverified first drafts, retrieval-augmented systems be preferred for regulatory research, client-specific information not be submitted to consumer LLMs, all citations be independently verified, and professional judgment remain the basis for advice and opinions.

Introduction

The adoption of Large Language Models (LLMs) — AI systems capable of generating human-like text responses to natural language queries — has accelerated significantly across professional services, including tax and compliance practice. Tools such as ChatGPT, Claude, Gemini, and their derivatives are increasingly being used by practitioners to research provisions, draft opinions, summarise judgments, and respond to client queries.

This article examines the specific risks that LLMs present in a tax and compliance context, with particular reference to Indian regulatory practice, and proposes a framework for their appropriate use consistent with professional obligations under the Chartered Accountants Act, 1949 and applicable standards of the Institute of Chartered Accountants of India (ICAI).

I. The Hallucination Problem and Its Implications for Tax Practice

What Hallucination Means

Hallucination, in the context of LLMs, refers to the generation of factually incorrect information presented with the same confidence and structure as correct information. This is not an occasional malfunction — it is a structural characteristic of how these models generate text. LLMs predict statistically likely text sequences based on their training data; they do not verify facts, cross-reference sources, or distinguish between knowledge and generation.

Practical Implications for Tax Practitioners

A practitioner researching the applicability of Section 194Q of the Income Tax Act, 1961 (now corresponding provisions under the Income Tax Act, 2025) using a consumer LLM may receive a detailed, structured response citing specific thresholds, applicability conditions, and rate structures. That response may reflect provisions as they stood at the model’s training cutoff — prior to subsequent amendments, CBDT circulars, or Finance Act changes — without any indication that the information may be superseded.

Similarly, practitioners researching GST provisions — particularly in areas subject to frequent council decisions and circulars, such as classification of services, input tax credit eligibility, or place of supply determinations — face the risk of receiving responses that do not reflect the current position as clarified by the GST Council or competent authority.

The professional risk is significant. A tax opinion or compliance position based on superseded law or an incorrectly stated provision is not merely incorrect — it may constitute professional negligence if the practitioner has failed to independently verify the LLM output against primary sources.

II. Temporal Blindness: The Training Cutoff Problem

LLMs are trained on data available up to a specific date, commonly referred to as the training cutoff. After this date, the model has no knowledge of developments in law, regulation, or judicial interpretation.

For Indian tax practitioners, this presents particular challenges:

Legislative changes: The Income Tax Act, 2025, which replaced the Income Tax Act, 1961 with effect from April 1, 2026, represents a significant legislative change that models trained before its enactment will not reflect. Practitioners using LLMs to research income tax provisions must be aware that responses may be based entirely on the superseded Act.

CBDT circulars and notifications: The Central Board of Direct Taxes issues circulars and notifications on a regular basis that clarify, modify, or withdraw earlier positions. LLMs with older training cutoffs will not reflect these developments.

GST Council decisions: GST law in India has been subject to significant evolution through council decisions, notifications, and circulars. A model’s knowledge of GST provisions reflects the state of law at its training cutoff and may not incorporate subsequent clarifications.

Judicial pronouncements: Significant judgments of the Supreme Court and various High Courts that alter the interpretation of tax provisions post the training cutoff will not be reflected in LLM responses.

The practical implication is that every LLM response on a tax or regulatory matter must be treated as potentially outdated and verified against current primary sources before being relied upon in any professional capacity.

III. The Absence of Chain of Custody

A defensible tax or compliance position requires a verifiable chain of reasoning from facts to statutory provision to interpretation to conclusion. Each step in that chain must be traceable to a specific, verifiable source.

LLMs provide conclusions. They do not provide chains of custody.

Even when an LLM response includes section references, circular numbers, or judgment citations, those references cannot be assumed to be accurate. A well-documented failure mode of LLMs is the generation of plausible-looking but fictitious citations — section numbers that do not exist, circular dates that do not correspond to the described content, judgment references attributed to the wrong court or date. This phenomenon, sometimes called “citation hallucination,” is particularly dangerous in legal and tax contexts because a fictitious citation is indistinguishable from a real one without independent verification.

For ICAI members, this raises a specific concern. The Standards on Auditing and the Code of Ethics require that professional opinions and conclusions be supported by sufficient appropriate evidence. An LLM output — even a correct one — does not constitute evidence in this sense. It constitutes a starting point for research that must be independently verified before being relied upon.

IV. Data Privacy and Professional Confidentiality

Regulatory Incidents

Several documented incidents are relevant to practitioners considering the use of consumer LLM tools for professional work:

In March 2023, a technical bug in OpenAI’s systems resulted in some users being able to view portions of other users’ conversation histories, including personal and payment information, before the issue was identified and resolved.

Samsung Electronics, following an internal incident in which engineers submitted proprietary source code and internal meeting notes to ChatGPT, implemented a company-wide ban on the use of external AI tools for work purposes, citing data security concerns.

The Italian data protection authority (Garante) temporarily suspended ChatGPT’s operation in Italy in March 2023, citing the absence of adequate legal basis for the processing of personal data under the General Data Protection Regulation (GDPR) and insufficient age verification mechanisms.

Multiple major financial institutions — including JPMorgan Chase, Goldman Sachs, and Citigroup — have implemented restrictions on employee use of consumer LLM tools, citing data security and confidentiality concerns.

Implications Under ICAI’s Code of Ethics

The ICAI Code of Ethics imposes on members a fundamental principle of confidentiality, requiring that members refrain from disclosing confidential client information to third parties without appropriate authority.

The submission of client-specific information — names, transaction details, financial positions, legal strategies under consideration — to a consumer LLM raises serious questions under this principle. The terms of service of most consumer LLM providers permit the use of submitted data for model improvement purposes, though opt-out mechanisms exist in varying forms. The data security practices of these providers, while generally robust, have demonstrated vulnerabilities.

Members are advised to exercise extreme caution before submitting any client-specific information to consumer LLM tools and to consider whether such submission is consistent with their confidentiality obligations under the Code of Ethics.

V. The Underrepresentation of Indian Law in LLM Training Data

A structural limitation of significant relevance to Indian practitioners is the underrepresentation of Indian legal and regulatory material in LLM training datasets relative to US and UK legal content.

Major LLMs have been trained predominantly on English-language content available on the internet, in which US and UK legal and regulatory material is substantially more prevalent than Indian material. The consequence is that these models tend to have more reliable and detailed coverage of US tax law, UK company law, and similar bodies of law than of Indian income tax provisions, GST regulations, SEBI rules, or ICAI standards.

This asymmetry is not visible to the user. An LLM response on an Indian GST provision may appear as confident and well-structured as its response on a US federal tax matter, even if the underlying training data for the Indian provision is significantly thinner. Practitioners should not assume that the quality of LLM responses on Indian regulatory matters is comparable to its quality on matters where training data is more robust.

VI. A Framework for Appropriate Use

The foregoing analysis does not support a blanket prohibition on the use of LLMs in tax and compliance practice. These tools offer genuine utility for general research, drafting, and the exploration of unfamiliar areas of law. The appropriate response is a disciplined framework for their use that accounts for their specific limitations.

1. Treat all LLM outputs as unverified first drafts. No LLM output should be relied upon in a professional context without independent verification against current primary sources — the bare Act, applicable rules, current CBDT circulars, GST notifications, or other directly applicable authority.

2. Prefer retrieval-augmented systems for regulatory research. Where AI assistance is used for regulatory research, practitioners should prefer systems that retrieve answers from curated, versioned databases of regulatory sources rather than relying on the model’s parametric knowledge. Such systems can provide citations to specific documents and versions, enabling verification.

3. Never submit client-specific information to consumer LLM tools. Client names, transaction details, financial positions, and legal strategies should not be submitted to consumer LLM tools. Where AI assistance is required for client-specific work, practitioners should use tools that offer adequate data isolation and do not use submitted content for model training.

4. Verify all citations independently. Every section reference, circular number, notification date, and judgment citation in an LLM response should be independently verified against the original source before being relied upon or communicated to a client.

5. Apply professional judgment to all outputs. LLM outputs should be treated as a research starting point, not a conclusion. The professional judgment of the practitioner — informed by current, verified knowledge of the applicable law — remains the basis for any opinion, position, or advice.

Conclusion

Large Language Models represent a significant development in professional research tools. Their ability to synthesise information, structure responses, and explore unfamiliar areas quickly offers genuine utility to tax and compliance practitioners.

However, their specific failure modes — hallucination, temporal blindness, the absence of verifiable chain of custody, data privacy risks, and the underrepresentation of Indian law in training data — present risks that are particularly consequential in a compliance context.

The practitioner who uses these tools with an accurate understanding of their limitations, applies independent verification to all outputs, and exercises professional judgment throughout is well-positioned to benefit from them. The practitioner who treats LLM outputs as authoritative, cites them without verification, or submits client-specific information without considering confidentiality obligations is exposed to significant professional and legal risk.

In tax and compliance practice, an answer without a verifiable evidentiary basis is not an answer. It is a liability.

Join Taxguru’s Network for Latest updates on Income Tax, GST, Company Law, Corporate Laws and other related subjects.

Leave a Comment

Your email address will not be published. Required fields are marked *

Search Post by Date
July 2026
M T W T F S S
 12345
6789101112
13141516171819
20212223242526
2728293031