A new wave of scrutiny is hitting the consulting industry’s “thought leadership” machine, and this time the spotlight is on how AI is being used to produce public-facing reports that are supposed to reflect the highest standards of rigor.
According to a report published by the Financial Times, PwC—one of the world’s Big Four professional services firms—has released some thought leadership materials that were found to contain errors consistent with AI hallucinations. In other words, parts of the published work appear to include statements that read like facts but are not reliably grounded in verifiable sources. The concern is not simply that the writing was sloppy or that arguments were weak; it’s that the content may have been generated or shaped in ways that allowed incorrect information to slip through while still being presented with confidence.
This is a particularly uncomfortable story for a firm that has positioned itself as an authority on AI and advanced analytics. PwC has spent years marketing its ability to help clients adopt AI responsibly, design governance frameworks, and translate complex technologies into business value. When a firm that sells expertise in AI also publishes materials that appear to contain AI-like errors, the reputational risk becomes more than a PR problem—it becomes a credibility problem.
What makes the issue stand out is the nature of thought leadership itself. These reports are not internal memos or drafts shared with a small group. They are meant to be persuasive, authoritative, and durable. They often influence how executives understand emerging risks, how boards frame strategy, and how organizations decide what to invest in. Thought leadership is supposed to set a benchmark: if a firm can’t get the basics right in a public report, what does that imply about the quality control behind the scenes?
The Financial Times report suggests that the problem wasn’t only about tone or framing. It points to the possibility that some of the inaccuracies resemble the classic failure mode of generative AI systems: hallucination. Hallucinations occur when an AI model produces plausible-sounding text that is not supported by reliable evidence. The output can look coherent, even sophisticated, and still be wrong—sometimes in subtle ways that are hard to detect without careful verification.
That distinction matters. Many readers assume that AI-generated content is either obviously wrong or obviously fabricated. But hallucinations don’t always announce themselves. They can appear as incorrect statistics, misattributed claims, references to studies that don’t exist, or descriptions of technical capabilities that don’t match reality. In a thought leadership context, where the writing is designed to sound confident and forward-looking, these errors can blend into the narrative rather than disrupt it.
So how does this happen in practice? The answer is likely less dramatic than people imagine, and more systemic.
In many organizations, AI tools are now embedded in workflows for drafting, summarizing, and brainstorming. A common pattern is that staff use AI to accelerate early-stage writing: generating outlines, producing first drafts, rephrasing complex concepts, or suggesting “insights” based on prompts. Even when teams intend to verify everything, the speed advantage can create a dangerous mismatch between production timelines and review capacity.
If a report is due quickly, the temptation is to treat AI output as a starting point rather than a claim that must be independently validated line by line. That’s where hallucinations become a risk. The model doesn’t “know” what it doesn’t know; it generates text that fits the style and structure requested. If the human reviewer focuses on readability, coherence, and overall argument rather than source-level accuracy, incorrect details can survive.
There’s also a second, quieter mechanism: the way AI can shape emphasis. Even if the core facts are correct, AI can nudge the narrative toward certain conclusions by selecting which details to highlight and which to omit. That can lead to a report that feels insightful but subtly misrepresents the evidence base. In other words, the problem may not always be that the report contains outright falsehoods; it may also be that it presents a skewed version of reality—one that is difficult to challenge because it is written in a persuasive voice.
For PwC, the stakes are amplified by the firm’s market position. Big Four firms operate at the intersection of advisory services, audit-related credibility, and technology consulting. Their brand depends on trust. Clients expect that when these firms speak publicly about AI governance, risk management, or regulatory trends, they are drawing from both expertise and disciplined verification.
When the Financial Times report describes “slapdash” work, it implies that the quality assurance process may not have been robust enough for the specific risks posed by AI-assisted drafting. That doesn’t necessarily mean that every part of the report was generated by AI. It could mean that AI was used in ways that reduced the friction of writing—without matching that reduction with equally strong verification steps.
This is where the broader industry lesson emerges.
AI hallucinations are not just a technical curiosity; they are a governance problem. They force organizations to rethink what “review” means. Traditional editorial review checks grammar, structure, and logic. But hallucination risk requires a different kind of diligence: verifying claims, validating citations, checking numbers, and ensuring that any referenced research actually exists and supports the statement being made.
In other words, the standard for “good writing” is not the same as the standard for “reliable writing.”
And reliability is exactly what thought leadership is supposed to deliver.
The unique twist in this story is that it comes from a firm that has been selling AI expertise. That creates a paradox: the same technology that can help produce faster analysis can also undermine the credibility of the analysis if it isn’t controlled properly. For executives and boards, this is not a theoretical concern. It affects procurement decisions, vendor selection, and the internal policies organizations adopt when they start using AI in their own communications.
If a major consultancy can publish materials with AI-like errors, then the question becomes: how many other organizations are doing something similar, perhaps with less visibility?
It’s also worth noting that thought leadership is often produced under commercial incentives. Firms want to demonstrate relevance. They want to be seen as early movers. They want to publish frequently enough to stay top-of-mind. Those incentives can unintentionally reward speed over verification, especially when AI tools make drafting faster and cheaper.
This is not an argument against AI. It’s an argument for aligning incentives and controls with the realities of generative systems.
One of the most important shifts organizations need is to treat AI output as untrusted until verified. That sounds obvious, but in practice it’s easy to drift. Teams may begin by verifying key claims, then gradually reduce the depth of verification because the output “looks right.” Over time, the organization develops a false sense of safety based on fluency.
Fluency is not evidence.
Generative AI can produce text that reads like it belongs in a high-quality report. It can mimic academic tone, corporate language, and even the cadence of credible analysis. But the model’s ability to sound authoritative is not the same as its ability to be accurate.
This is why the most effective governance approaches tend to be procedural rather than purely cultural. Culture matters, but procedures reduce the chance of human error under time pressure.
What might stronger procedures look like in a thought leadership workflow?
First, there should be a clear separation between drafting and claiming. AI can draft, summarize, and propose structures. But when a report includes factual assertions—statistics, regulatory interpretations, study findings, or technical descriptions—those claims should be traceable to sources that can be checked independently. If a report cites research, the cited research should be verified. If it includes numbers, those numbers should be reproducible from the underlying data or documented sources.
Second, teams should implement “citation discipline.” Many hallucinations involve references that are fabricated or misrepresented. A citation discipline approach ensures that every citation is real, accessible, and relevant to the claim it supports. This is especially important for reports that aim to influence policy or executive decision-making.
Third, organizations should define what counts as a “high-risk” claim. Not all statements require the same level of verification. But claims about legal compliance, safety, performance metrics, or regulatory requirements should be treated as high-risk. Those are the areas where errors can cause real harm—financially, legally, or operationally.
Fourth, there should be a review step that is designed specifically to catch AI-style errors. Traditional editing may not be enough. A specialized review can focus on factual consistency, source validation, and internal coherence between claims and evidence. Some teams use checklists; others use tooling that flags suspicious citations or inconsistencies. The key is that the review must be built for the failure modes of generative AI.
Fifth, organizations should track where AI was used. If a report includes AI-assisted drafting, the workflow should record that fact so that reviewers know where to focus. Transparency inside the organization helps ensure that verification effort is targeted rather than distributed evenly across the document.
None of these steps guarantee perfection. But they reduce the probability that hallucinations will survive into publication.
The PwC case, as described by the Financial Times, also raises a more uncomfortable question: what happens when the public-facing narrative is built on a foundation that wasn’t fully verified?
Thought leadership is often used by clients as a proxy for expertise. When a firm publishes a report, it signals that the firm stands behind the content. If the content contains errors, the damage extends beyond embarrassment. It can distort how clients interpret AI risks, how they prioritize governance measures, and how they communicate internally about what AI can and cannot do.
For example, if a report overstates what AI systems can reliably infer, or mischaracterizes how certain governance frameworks work, it can lead organizations to adopt policies that are either too lax or misdirected. The cost of misinformation is not limited to reputational harm; it can show up later as poor decisions, wasted budgets, or compliance gaps.
This is why the story resonates beyond the consulting industry. It’s a reminder that AI hallucinations are not confined to chatbots. They can enter the mainstream through professional workflows—especially when those workflows are optimized for speed
