How to avoid AI presenting misinformation about your brand
The AI era is changing the way people search for and consume information. Tools like ChatGPT and Google's AI Overviews mean people increasingly get answers from AI rather than a list of links from a Google search to go and visit. In some cases, this eliminates the need to visit websites altogether.
One result of this is a decline in website visits, which anecdotally many digital teams report. Currently many organisations are scrambling to meet this change in user habits, heavily investing in AEO, GEO and related practices in order to make their brand more visible in AI-powered summaries and responses.
But one area where there is arguably less awareness is in taking actions to reduce the chances of AI tools misrepresenting your brand and providing answers which are inaccurate. AI “misinformation” can mean AI tools like ChatGPT providing information about your company, products and services that is incorrect or out of date. The impact of this can vary from something that is limited but annoying – like store hours that are wrong – to something more serious where in a regulated industry inaccurate information could create compliance risks.
Of course, AI tools hallucinate and produce inaccuracies, but sometimes the reason for AI misinformation is actually in the control of the organisation. In this post we look at four causes of AI misinformation that brands can do something about.
1. Out-of-date content across your digital estate
LLMs and related AI tools surface content on the web to provide responses and answers. If your digital estate has content that is out of date or inaccurate then there is a chance that this will be reported faithfully back as fact.
Digital estates tend to grow and proliferate over the years, sometimes extending across multiple sites which then get forgotten about. Remember that microsite that was set up for a campaign a few years back and is still active? Teams also inherit long-forgotten sites when they acquire other businesses. The main corporate website itself can also be subject to sprawl. Research from AAANow.AI suggests that as much as 41% of a digital estate might not even be known to the digital team.
Cleaning up or archiving old content and sites is often at the bottom of the “to do” list for busy digital marketing teams. Rather than being an ongoing activity as part of regular website management, it becomes a clean-up project in its own right which then gets deprioritised because there are always other things to do. While people might seldom visit this content, AI will access it all and use it to formulate responses, and not necessarily know years-old information is either incorrect or has been superseded.
2. Content is not structured for AI to interpret
We’ve previously covered some of the things you can do to make content more interpretable for AI, so it can identify key facts about your organisation and your products. Approaches include adding schema markup which indicates what your content represents, for example the core information about your business. Other tactics include having clear headings and structured content such as tables and FAQs. Additionally, Microsoft warns about hiding important content, for example in tabs or expandable menus. The point here is to ensure that content is optimised so the chances of content being missed or misinterpreted are reduced.
3. AI tools are being blocked
One potential reason LLMs and related AI tools are more likely to provide misinformation about your organisation is that they are blocked from surfacing your authoritative and current content, meaning that they may turn to less authoritative content to provide responses, potentially from third-party sources outside your control.
AI tools get blocked deliberately but also inadvertently. There are legitimate reasons why some AI tools may effectively be blocked from accessing your content. For example, this could be for security or performance issues, or it may be a commercial decision because content is licensed. It might also be to prevent intellectual property (IP) being used to train AI models.
But sometimes measures that may be in place might accidentally block access. These include CDN or security settings which are just a little too stringent, or rules set up to stop bots scraping your content, but are now preventing AI crawlers from accessing content. There may also be errors in the setup, for example with robots.txt.
Given the elevated importance of AI responses, it is worth checking or reviewing to see if there are any AI tools blocked. Note that it's also possible to block bots that collect training data while still allowing those that power AI search, so IP policies can be maintained without reducing visibility or increasing the risk of misinformation.
4. PDFs are not properly configured for AI
PDFs remain a problem area for digital management and AI misinformation. Often, they are authoritative, official content – for example a publication, a brochure, a key policy, a user guide, official guidance or an annual report; this is the kind of content that AI should surface to give the right answers. But PDFs also tend not to get updated, either because the content is more expensive to update compared to a page, or because it needs to remain available as it was issued for regulatory or compliance reasons. This means PDFs accessible via a website might not always have up-to-date information.
But there can also be issues with PDFs that might all look fine for human readers, but actually prevent AI interpreting information properly. For example, a PDF without accessibility tags loses its underlying structure, so AI tools may misread the order of content or pull figures from the wrong row or column of a table, leading to a greater chance of misinformation. Scanned PDFs can also be an issue, as they may contain no machine-readable text.
Ensuring PDFs are correctly tagged is key. You can also ensure superseded PDFs are excluded from search indexes, using a “noindex” instruction which should reduce the chance of them being cited in AI responses.
Avoiding AI misinformation
As more people rely on AI-generated answers in searching, we think issues around AI misinformation will grow in importance. While this isn’t entirely in an organisation’s control, digital teams can reduce the risks by removing redundant content, ensuring AI tools can access key content, and structuring content to make it easily read by AI.
If you’d like to discuss how to avoid AI misinformation on your website, then get in touch.
Related Blog posts
Unlimited possibilities
3chillies
Get in touch today