The Australian government paid $440,000 AUD for Deloitte to analyse an aspect of its welfare system. What it got back was a report featuring made-up citations.
Why? Reports show that Deloitte Australia made use of OpenAI’s GPT-4o as part of its work, and it did what generative models do: it generated things, or to put it another way, made up AI slop.
These so-called hallucinations were spotted by an academic who noticed some of the papers didn’t actually exist, though some were attributed to a real professor. Another quote – from a federal justice – appears to also have been made up.
Deloitte has since updated its report, with 14 citations disappearing, and said it will repay the last segment of its contract, though it’s not clear how much money that will involve.
What’s the problem?
This is alarming on several fronts. First, Deloitte’s staff didn’t fact check their own report, despite the serious nature of the topic, which included automated financial penalties in the Australian welfare system directed at jobseekers. Second, it used AI without initially referencing it, meaning the methodology of the report wasn’t really honest or accurate. And third, you have to wonder why Deloitte used AI at all when the consultancy was being paid for its expertise and analysis.
The new version of the report now admits that a large language model was used for parts of the previous version. “The updates made in no way impact or affect the substantive content, findings and recommendations in the report,” Deloitte stated in the amended version, according to the FT.
The Australian government agreed and said the core of the recommendations of the review were not changed.
But many feel that’s not good enough. The fake citations were spotted by Sydney University Deputy Director of Health Law Chris Rudge, who reportedly told local newspapers that: “you cannot trust the recommendations when the very foundation of the report is built on a flawed, originally undisclosed, and non-expert methodology… Deloitte has admitted to using generative AI for a core analytical task; but it failed to disclose this in the first place.”
Deloitte still doubling down on AI
Despite that damning news story from down under, Deloitte this week announced a partnership with OpenAI rival Anthropic to roll out its Claude system across 470,000 staff and clients – perhaps the only lesson learned was to find a better LLM.
“As part of the collaboration, Deloitte will establish a Claude Center of Excellence with trained specialists who will develop implementation frameworks, share leading practices across deployments, and provide ongoing technical support to create the systems needed to move AI pilots to production at scale,” Anthropic said in a statement.
Perhaps Deloitte might want to first brush up on its own understanding and use of AI before telling others how to go about deploying the technology.
Here’s some tips to get started: don’t hide how you use AI; fact check everything; and know that it makes things up. Otherwise it could cost you – and your clients – a contract and dent your reputation.
Related articles