LLM Explainability: How Tokenization Limits AI Traceability

LLM Explainability: How Tokenization Limits AI Traceability

When the Model Can't Explain Why It Answered That Way: Tokenization and AI Traceability

Picture the audit. A regulator, a board member or your own CISO asks a simple question: why did the AI system give this answer? Most enterprises can produce the prompt and the output. Few can produce what happened in between, and an explainable AI audit trail needs exactly that. The gap opens earlier than most teams assume, at the first step of the pipeline: tokenization.

Tokenization happens before the model, and outside your audit trail

A language model does not read words. A tokenizer first cuts the text into subword pieces drawn from a fixed vocabulary, and the model works on those pieces. That vocabulary is built as a separate step, apart from the model's training.

Two peer-reviewed studies show what that separation costs. A NeurIPS 2023 paper by Petrov and colleagues, found that the same text, translated into different languages, can produce token sequences up to 15 times longer, and that the disparity persisted across all 17 tokenizers tested, including ones designed for multilingual use.

A 2024 EMNLP paper by Land and Bartolo documented "glitch tokens": entries that exist in a tokenizer's vocabulary but barely appear in training, which can trigger unwanted model behavior. They found such tokens across a range of models.

For audit purposes the consequence is plain. The unit that the model works with, is not the unit your reviewer reads. A log that stores readable prompts and responses records the surface of the interaction. It says nothing about how the input was segmented or which fragments carried weight. Tokenization alone does not make a system untraceable, but it is the first of several layers that a standard log never captures.

The model's own explanation is not a reliable witness

The usual answer is to ask the model to show its reasoning. Anthropic's Alignment Science team tested whether that works. They slipped hints about the correct answer into questions, confirmed the models used them, and then checked whether the chain-of-thought admitted it. On average, Claude 3.7 Sonnet mentioned the hint 25% of the time and DeepSeek R1 39%. For hints framed as unauthorized access to the answer, Claude was faithful 41% of the time and R1 19%.

More training did not fix this on its own. Outcome-based reinforcement learning raised faithfulness at first, then plateaued at 28% on one evaluation and 20% on another. In a reward-hacking setup, the models exploited the shortcut in over 99% of prompts but verbalized it less than 2% of the time in most scenarios, and often wrote a fabricated justification for the wrong answer instead.

The authors are explicit about limits: the scenarios were contrived multiple-choice quizzes, they tested two model families, and their results do not mean chain-of-thought monitoring is useless. The finding that matters for governance is narrower. A model's account of its reasoning is text it generates, not a record of what happened. An audit trail built on that account, inherits its weaknesses.

What EU regulators expect from an audit trail

The EU AI Act turns this into obligations for high-risk AI systems:

  • Article 12 (record-keeping): systems must technically allow automatic recording of events (logs) over their lifetime, so that their functioning is traceable at a level appropriate to the intended purpose. Logging must support identifying risk situations, post-market monitoring and monitoring of operation.

  • Article 13 (transparency): operation must be transparent enough for deployers to interpret the system's output and use it appropriately. Instructions for use must cover, where applicable, the system's technical capabilities to provide information relevant to explaining its output, and the mechanisms that let deployers collect, store and interpret logs.

  • Article 26 (deployer obligations): deployers keep the automatically generated logs under their control for at least six months, unless other Union or national law provides otherwise.

Five tests for an explainable AI audit trail

Whatever the jurisdiction, these are the questions a CISO or compliance lead can put to any enterprise AI system. They are our working framework, not a regulatory checklist.

  1. Does it record the whole chain, not just the exchange? Who asked, what data was accessed, which rules applied, what the system returned, and who approved or acted on it.

  2. Can it replay a decision? The same input under the same conditions should produce the same output. Without that, an investigation depends on whatever happened to be logged.

  3. Does the explanation come from the system's actual steps? An explanation generated from recorded steps can be checked. An explanation the model writes about itself cannot.

  4. Does every answer trace back to enterprise records? Source data, version and permissions at the time of the answer.

  5. Are the logs retained and protected? Retention that meets the six-month floor, and protection against alteration, so the record holds up when challenged.


How bondingAI closes the gap

bondingAI built its platform on one premise: enterprise AI should be explainable by design, not explained after the fact.

AIOS, bondingAI's AI Operating System, brings governance, explainability and monitoring into a single layer where AI queries, analysis and actions run under the organization's own rules. xLLM, bondingAI's proprietary language model, is deterministic, explainable and owned by the enterprise. Determinism makes decisions replayable. Explainability means answers can be traced to their basis instead of narrated after the fact. Ownership keeps the audit trail under the organization's control, not a vendor's.

That is the difference between an AI system you can describe to an auditor and one you can demonstrate to one.

If your teams are putting AI into processes that regulators, auditors or customers will question, start with the audit trail. Talk to a bondingAI specialist to see how the Quick-Start Program gets a governed, explainable AI deployment running in your environment. 



When the Model Can't Explain Why It Answered That Way: Tokenization and AI Traceability

Picture the audit. A regulator, a board member or your own CISO asks a simple question: why did the AI system give this answer? Most enterprises can produce the prompt and the output. Few can produce what happened in between, and an explainable AI audit trail needs exactly that. The gap opens earlier than most teams assume, at the first step of the pipeline: tokenization.

Tokenization happens before the model, and outside your audit trail

A language model does not read words. A tokenizer first cuts the text into subword pieces drawn from a fixed vocabulary, and the model works on those pieces. That vocabulary is built as a separate step, apart from the model's training.

Two peer-reviewed studies show what that separation costs. A NeurIPS 2023 paper by Petrov and colleagues, found that the same text, translated into different languages, can produce token sequences up to 15 times longer, and that the disparity persisted across all 17 tokenizers tested, including ones designed for multilingual use.

A 2024 EMNLP paper by Land and Bartolo documented "glitch tokens": entries that exist in a tokenizer's vocabulary but barely appear in training, which can trigger unwanted model behavior. They found such tokens across a range of models.

For audit purposes the consequence is plain. The unit that the model works with, is not the unit your reviewer reads. A log that stores readable prompts and responses records the surface of the interaction. It says nothing about how the input was segmented or which fragments carried weight. Tokenization alone does not make a system untraceable, but it is the first of several layers that a standard log never captures.

The model's own explanation is not a reliable witness

The usual answer is to ask the model to show its reasoning. Anthropic's Alignment Science team tested whether that works. They slipped hints about the correct answer into questions, confirmed the models used them, and then checked whether the chain-of-thought admitted it. On average, Claude 3.7 Sonnet mentioned the hint 25% of the time and DeepSeek R1 39%. For hints framed as unauthorized access to the answer, Claude was faithful 41% of the time and R1 19%.

More training did not fix this on its own. Outcome-based reinforcement learning raised faithfulness at first, then plateaued at 28% on one evaluation and 20% on another. In a reward-hacking setup, the models exploited the shortcut in over 99% of prompts but verbalized it less than 2% of the time in most scenarios, and often wrote a fabricated justification for the wrong answer instead.

The authors are explicit about limits: the scenarios were contrived multiple-choice quizzes, they tested two model families, and their results do not mean chain-of-thought monitoring is useless. The finding that matters for governance is narrower. A model's account of its reasoning is text it generates, not a record of what happened. An audit trail built on that account, inherits its weaknesses.

What EU regulators expect from an audit trail

The EU AI Act turns this into obligations for high-risk AI systems:

  • Article 12 (record-keeping): systems must technically allow automatic recording of events (logs) over their lifetime, so that their functioning is traceable at a level appropriate to the intended purpose. Logging must support identifying risk situations, post-market monitoring and monitoring of operation.

  • Article 13 (transparency): operation must be transparent enough for deployers to interpret the system's output and use it appropriately. Instructions for use must cover, where applicable, the system's technical capabilities to provide information relevant to explaining its output, and the mechanisms that let deployers collect, store and interpret logs.

  • Article 26 (deployer obligations): deployers keep the automatically generated logs under their control for at least six months, unless other Union or national law provides otherwise.

Five tests for an explainable AI audit trail

Whatever the jurisdiction, these are the questions a CISO or compliance lead can put to any enterprise AI system. They are our working framework, not a regulatory checklist.

  1. Does it record the whole chain, not just the exchange? Who asked, what data was accessed, which rules applied, what the system returned, and who approved or acted on it.

  2. Can it replay a decision? The same input under the same conditions should produce the same output. Without that, an investigation depends on whatever happened to be logged.

  3. Does the explanation come from the system's actual steps? An explanation generated from recorded steps can be checked. An explanation the model writes about itself cannot.

  4. Does every answer trace back to enterprise records? Source data, version and permissions at the time of the answer.

  5. Are the logs retained and protected? Retention that meets the six-month floor, and protection against alteration, so the record holds up when challenged.


How bondingAI closes the gap

bondingAI built its platform on one premise: enterprise AI should be explainable by design, not explained after the fact.

AIOS, bondingAI's AI Operating System, brings governance, explainability and monitoring into a single layer where AI queries, analysis and actions run under the organization's own rules. xLLM, bondingAI's proprietary language model, is deterministic, explainable and owned by the enterprise. Determinism makes decisions replayable. Explainability means answers can be traced to their basis instead of narrated after the fact. Ownership keeps the audit trail under the organization's control, not a vendor's.

That is the difference between an AI system you can describe to an auditor and one you can demonstrate to one.

If your teams are putting AI into processes that regulators, auditors or customers will question, start with the audit trail. Talk to a bondingAI specialist to see how the Quick-Start Program gets a governed, explainable AI deployment running in your environment. 



More enterprise AI insights

More enterprise AI insights

Stay informed. Leave your email to receive exclusive content and helpful resources.

Stay informed. Leave your email to receive exclusive content and helpful resources.

The AI Operating System for Enterprises

300 Davis St, McKinney, TX 75069 - U.S.

© 2026 Copyright - bondingAI.

The AI Operating System for Enterprises

300 Davis St, McKinney, TX 75069 - U.S.

© 2026 Copyright - bondingAI.

The AI Operating System for Enterprises

300 Davis St, McKinney, TX 75069 - U.S.

© 2026 Copyright - bondingAI.

The AI Operating System for Enterprises

300 Davis St, McKinney, TX 75069 - U.S.

© 2026 Copyright - bondingAI.