Skip to Content
Article

Issue #177 | Claudits, Errors & Smart Humans

by Sam Tomlinson
July 19, 2026

If your inbox looks anything like mine, you’re getting pitched an AI audit tool once or twice a week. “Upload your account, find every issue & increase your ROAS 24% in 90 seconds” (yes, that’s an actual email subject line). I’ve rejected every single one – at least, up until now. But, while I was bored on vacation (and with a LOT of urging from a good friend), I finally caved. I ran one — a “Claudit” — on a live client account we’re currently auditing the old-fashioned way.

In full transparency: the audit in question is a paid audit. We were brought in to provide a second set of eyes on their Google + Meta Ads setup & provide recommendations. Unlike many “free” audits, this isn’t a sales pitch or agenda-driven audit – it’s just a project-based engagement where our only goal is to help the brand. And, to be even more blunt: this brand would not be a fit for us as a long-term engagement (this is something we’ve told them at the outset). The confluence of these things – me having some spare time, some gnawing curiosity, and a contemporaneously-created human audit – resulted in a rare opportunity to compare man vs. machine in a (relatively) unbiased, neutral manner.

The results were eye-opening. And the lessons learned from comparing the Claudit to our work are shaping how we integrate AI into our own audit/review processes going forward – and I hope sharing some of what I’ve seen helps you avoid the pitfalls + increase the value you’re receiving from AI Tools.

Before we dive into what we found, some context: we provided the Claudit tool the same information we received, completely anonymized. Everything was conducted in a secure environment. In addition to ad account structure & performance data, we also provided the same business intelligence (baseline AOV / contribution margin by service), approximate capacity by both client tier (enterprise, mid-market, SMB, entry/startup) and service line, along with relevant brand guidelines, operating procedures, limitations, etc. The platform itself was a thin Claude wrapper – a combination of the tool’s prompts + a similar-to-Claude interface that enabled interaction throughout the process.

Once we loaded all that up, the AI audit came back with a scathing indictment – 92 issues with the Google account and ~56 with Meta (depending on how you count) across ~30 total pages, with each issue seemingly more egregious than the last. My genuine first reaction when opening it was…”How did we miss all this?” – followed closely by “Holy [expletive] this is bad.” And if that was my reaction, I can only imagine how a SMB owner/CEO with limited marketing knowledge would react if they ran their accounts through this tool. The most likely reaction would be shock…followed closely by outright fury at their marketing team.

But, since this was vacation and I had nothing better to do, I started *really* reading the report – going issue-by-issue, line-by-line – to understand how we missed all of this. And the more I read, the more I found myself saying, “this isn’t right,” and “that doesn’t follow.” It turns out confidence is not the same as validity, and a long list of problems doesn’t always mean everything is wrong.

I started characterizing every claimed issue into a 2×2 matrix – with “legitimacy” on the x axis (i.e. the issue is real/legitimate) and “impact” (defined as how much the issue likely impacted results/account performance + how urgent the item fix was). As a few examples of how I scored these:

  • An issue cited in the AI audit was that branded search terms were leaching into non-branded campaigns; in reality, there were 3 clicks on mis-spelled brand/product terms that had made their way into a non-branded campaign. I scored this low legitimacy (3 clicks on an account spending $100k/mo is nothing) and low impact (excluding those clicks from that campaign would have had no impact on performance).
  • A second issue was that a high-performing non-branded campaign was limited by budget for ~20 of the last 30 days. I scored this high legitimacy and high impact (reallocating some budget to this campaign was likely to produce a material increase in leads).
  • A third issue was that multiple campaigns ran ad schedules preventing them from serving during specific overnight hours (11 pm to 5 am). This was scored as high legitimacy (the conclusion was right, and the off-hours coverage issue was real), low impact. Per the intake documentation, the client had ample CRM data showing that overnight inquiries tended to convert at much lower rates than extended-business-hours inquiries. Moreover, given the budget limitations (noted above), incremental dollars were likely better spent addressing the high-performing, limited by budget campaigns vs. spreading an already thin budget thinner.
  • A fourth issue was the campaign structure – the AI cited that separate non-branded and competitor campaigns was a structural miss, as “competitor” keywords are functionally non-branded. I scored this as low legitimacy and low impact – our data has consistently shown that (1) competitor terms perform best when segregated away from genuine non-branded terms (quality scores aggregate at the campaign level, so the 1’s / 2’s / 3’s that typically show up for competitor KWs drag down the 6’s/7’s/8’s we have in our non-branded campaigns) and (2) our best results from competitor campaigns have consistently come from highly targeted ads + tailored landers. The easiest way to manage that is by segregating those KWs into their own campaigns, with a single ad group per competitor.
  • A fifth/final issue was a claim that “Conversion Tracking Is Broken” because two events – a GA4 event (tracking demos booked as “qualified leads”) and the GTag + Offline Conversion event (tracking SQLs) – showed significantly different values. I categorized this as low legitimacy, high impact – the issue itself was an artifact of different events; only the GTag event was “primary” and used for optimization. If the recommendation had been followed (retagging everything), the result would have been quite bad for the account.

If you’re curious about how I characterized each quadrant (designed with Claude’s help)

As I categorized each issue, I noticed a major pattern: the overwhelming majority (~80%+) of the “issues” cited fell into the “low impact” hemisphere, with ~21% clustering into the low impact / low legitimacy quadrant. In other words, the AI was citing “legitimate” issues about 75% of the time, but the majority were nit-picking and/or irrelevant (like the misspelled branded search term leaking into a non-branded campaign, or a flag about pinned headlines in ads, or a flag that different campaigns had different tracking numbers).

In order to be fair, I did the same exercise with our own findings – and what I found was the polar opposite: 0 of the 22 recommendations we made were in the low legitimacy / low impact quadrant; 13 of the 22 were firmly in the “high impact / high legitimacy” quadrant – and of those 13, 6 were NOT found in the AI audit. That’s nearly half of the high-value findings that a client actually pays for and needs to see, completely omitted from the report.

Examples of those recommendations omitted from the AI audit:

  1. Landing Page structure / performance – we identified that LP performance was a major driver of conversion outcomes, but that the performance of the LPs was not correlated with the LP quality score component. The campaigns with the highest conversion rate were the ones using specific LP styles/structures – and we recommended building those out for the campaigns on the old templates. Despite being given the destination URLs, the AI never bothered to check the LPs.
  2. Audience Insight Gaps – we flagged that some non-obvious search terms / queries were driving disproportionately qualified + economical lead volume, but were being under-funded because they were phrase matched to more “generic” keywords. Our recommendation was to build those into their own ad group, so they’d get a dedicated budget.
  3. PMAX Performance – the AI raved about PMAX and recommended an increase in budget; we went the opposite way, precisely because we found that PMAX, as it was set up, was bidding on brand terms AND producing lower quality leads (as measured by lead-to-SQL rate). While the top-line lead numbers looked solid, when we dug into performance, we found that the drivers of that performance were better leveraged elsewhere.

As I evaluated both the distribution of the issues and the failure modes (the AI both recommended changes that were unlikely to move the needle AND missed opportunities that required reasoning or inference/judgment), I started to piece together 4 common threads running through the Claudit:

1. Everything Was an Emergency (So Nothing Was)

The AI audit provided next-to-no prioritization, even when I explicitly asked for it. A misspelled brand/product term leaking into a non-branded search campaign was given the same weight (and even the same number of words) as high-performing campaigns limited by budget, which got the same weight as ad copy, which got the same weight as change-history patterns. Each issue was treated like a 5-alarm fire, when the reality was far more nuanced.

The challenge with AI is that it has – effectively – endless capacity for content generation. And given unlimited resources, there is always unlimited work. Warren Buffett once said, “If a cop follows you for 500 miles, you’re gonna get a ticket” – that’s the same principle here. If I spend enough time parsing through an ad account, I’m going to find issues.

Which brings me back to the “everything’s an emergency” point – when everything is critical, nothing is. The single most valuable component of any audit is the ranking of problems/issues. What’s costing you the most money right now? What risks have the greatest potential to harm performance in the future? What looks bad based on “platform recommendations” or “best practices”, but is actually fine? What is just a stylistic preference vs. a legitimate gripe?

To be very clear, this account had real issues – but when the 5-10 real problems are interspersed alongside 130+ nothingburgers + housekeeping items, going from a damning report to actual progress is all but impossible.

2. The Claudit Smuggled In Assumptions It Never Disclosed

I’ll be the first to admit I believe in better starting points, not best practices. In my experience, many roads lead to Rome – and just because an account is set up in a way I don’t personally prefer does not automatically disqualify it. There are multiple valid ways to run any ad account (or any marketing program, for that matter) – which is what makes good audits so difficult: you must disintermediate your personal preferences and assumptions from the performance.

The Claudit was built on a set of unstated assumptions about how Google & Meta ad accounts should be structured and managed. Some of those assumptions were legitimate and defensible; some were not. But none were shared in the document.

Any honest evaluation of the road someone chose has to start by naming the map you’re using to evaluate it. Claude graded the account against an invisible, unstated ideal, then deemed every deviation a defect. That’s not an audit – it’s a sham that would make anyone look bad, simply because of the Buffett point above: if you look hard enough, for long enough (which AI can do), you’ll find a mistake in anyone’s work.

3. Selective Enforcement & Willful Ignorance

The above 2 issues are commonplace in AI contexts – I think most people who use Claude/GPT/Gemini in their day-to-day are familiar with their confident-sounding bullshittery and penchant for drama.

But I’ll be the first to admit I was surprised by this one. Claude flagged a non-branded search campaign for the client’s enterprise service because its CPA was 2.4x higher than the SMB non-branded campaign. The recommendation was to pause the enterprise campaign, as it was “wasteful” and “inefficient” (quoting exactly from the Claudit). It even included a small summary table that reinforced the point, with the SMB + enterprise CPAs compared and a nice, bolded conclusion.

The only problem: the conversion values were right there in the same view. Enterprise drove more than 6x the revenue of SMB (which, intuitively, makes sense – enterprise deals are much larger than SMB deals). So while it cost 2.4x more to get those deals, the ~6x higher profit per deal more than made up for the more expensive acquisition.

One of the most consistently problematic aspects of this audit was the presence of “landmine” recommendations like this one – confident-sounding, seemingly-data-supported recommendations that, if acted on, would materially harm the business. In the short-term, those are my greatest fear for most accounts (and any SMB or media buyer who trusts AI audits blindly) – yes, they’ll catch some legitimate issues, but all of the good/progress/savings from all of that can be detonated by implementing a single landmine recommendation. This is the same principle that Taleb argues for across the Incerto: the upside doesn’t matter if you can’t survive the cost of failure.

4. It Ignored Logical, Reasoned Decisions

One of the first things that it cited in the creative section was the substantial use of pinned headlines. Aside from the fact that I’m pretty pro-pinning (look at enough ads in the wild, and you too will become a believer), those headlines were pinned because a partnership agreement requires it. This was included in the intake/context document shared, but then ignored in the actual report.

The AI audit similarly flagged dayparting and lower weekend budgets as issues, when those were deliberate, tested decisions, backed up with CRM data showing that leads during those odd times convert at much lower rates + are much lower quality.

Finally, it accused the current team of “poor management” because most changes were made in bulk, once or twice a week. Bulk changes 1-2x per week, for an account of this size/scale/age, running primarily on tCPA, is prudent. If you’re pushing small, non-essential changes every day, you’re continuously disrupting learning. We’ve seen (and this is the rare case where Google also agrees) this results in higher CPAs, simply because the platform can never reach an equilibrium. And moreover, when a real issue arose – for instance, we observed a phrase match KW started matching to competitor names out of the blue – changes weren’t made immediately. So, this wasn’t a “set it and forget it” situation – the current management team deviated from their 1-2x changes/week SOP and added most of the competitor names as negatives the same day they appeared in the STR.

The common thread running through all of these examples is that the AI auditor identified a pattern it didn’t understand or didn’t align with its assumptions, then immediately jumped to the most damning explanation. On one hand, that makes some sense, given that these systems are trained on the internet (and anyone who has spent significant time on the internet is likely familiar with this pattern of behavior). But, on the other hand, it makes them very poor evaluators of nuanced performance. 90% of the value of an audit is what it doesn’t say – it’s the values of judgement, prudence, grace and empathy that go into separating the proverbial wheat from the chaff.

Claudits Are Here To Stay

The other thing that struck me in reading this audit was how easy it was to generate – it took less than 10 minutes to upload the relevant docs, provide the download of the account performance and link everything up. Then it was another ~15 minutes for the actual audit to run. In less than a How I Met Your Mother rerun, I had the full Claudit.

Compare that to our team’s audit of the same accounts, which took ~20 hours to put together (not including meetings, design time, editing, etc.), and it’s easy to see the appeal. Why pay someone thousands of dollars and wait weeks, when Claude can do it in less than an hour?

It’s for precisely those reasons – time and cost – that I fear Claudits are here to stay. They’re simply too easy – and when you layer on top of that ease of use the illusion of authority + the persuasive nature of how it presents information, the end result is a headache (and a mistake) waiting to happen.

In full disclosure, prior to this exercise, we had a few clients run Claudits on their accounts. Most were – to be candid and self-congratulatory at the same time – not nearly as damning as this one, but they still all found issues. We had good answers for most, and there were a few small things they found that were helpful. As Bruce Lee famously said, “Absorb what is useful, discard the rest.”

How We’re Actually Integrating AI Into Our Audits

If you’ve made it this far, you’re probably thinking, “There’s no way in hell Sam trusts AI tools to do audits.” – but (plot twist) that’s not the case at all. I’m not telling you to avoid these tools. I’m telling you to put them in the right seat.

Here’s how we’re using AI & what processes we’re putting in place to reap the benefits it can produce, while minimizing the (very bad) side effects:

Use AI To Cast A Wide Net, But Not Make Judgments

Claude (and ChatGPT / Gemini) is genuinely good at surfacing things a human might miss, such as the misspelled brand term leaking into a non-brand campaign, or the negative keyword that’s blocking a valid search term, or the two RSAs running in the same ad group with near-identical headlines, or the two slightly-different variants of a YouTube audience, or the campaigns that are supposed to be covering the same geos, but one has a few extra zip codes included. Each one of those is a legitimate issue, even if all of them are relatively low-impact.

Then, once you have the list of issues, have a real person help sort them into priority tiers – the machine creates the list of items to check, the person validates the issues (especially the high-impact ones) and sorts them into a list.

End result, thus far: 20%-30% reduction in time to review an account across the 3 small audits/reviews we’ve started since. This is far from definitive data, but it’s a very promising start.

Let It Handle Rote Tasks

There are dozens of tasks in any audit that are simultaneously incredibly time consuming and have high degrees of subjectivity – think: evaluating RSA copy; assessing lead quality; determining review sentiment; evaluating creative diversity/depth, reviewing blogs/articles, etc. We do them as part of an audit, but – being completely honest – there’s massive subjectivity in each one, especially if more than 1 person is involved or the task itself is substantial (3+ hours of work). If you’ve ever tried to read and assess 1,000+ reviews, you know that your definition of “good” or “bad” or “neutral” likely slips hour-by-hour, day-by-day.

These are EXACTLY the types of tasks where AI has been wonderfully helpful – especially when we’ve given it an explicit evaluation rubric AND a charge to grade probabilistically vs. deterministically. The net-net of that is that, instead of the machine giving a confident response, it spits out a “This ad scored a 7/10 for differentiation and a 3/10 for originality. It uses the same structure as 17 other ads in the account.”

That is exponentially more helpful than a generic conclusion – precisely because it provides our team with the information needed to make a diagnosis, without prejudging that conclusion. End result is a better, more consistent result for our client – and our team has the confidence in using it, because AI brings a level of stability and consistency to the evaluation that is difficult for a person to match.

Force The AI To Show Assumptions

We now explicitly include language in our audit prompts that force the AI to disclose the assumptions it is using to evaluate the account and surface issues, before it lists what it found. We also include specific language (such as the below) intended to force a question back vs. a verdict on incomplete information.

The related instruction that’s been helpful: when the AI makes a recommendation or assumption, force it to cite a source for the assumption and a justification.

Here’s the initial prompt I used (but feel free to adjust):

“In your response, indicate each assumption made or best practice considered, the source of that assumption or best practice, and under what conditions or situations this recommendation would not make sense.”

Then, Have It Project Results From Recommended Changes

I’ve also started forcing the AI to project the possible impacts from recommended changes on the account, along with a probability of each outcome. That may sound odd, but it quickly exposes situations where either: (a) the impact of the change is small to non-existent or (b) the change has a modest probability of producing a substantial negative impact on the account or business. For example, when I went back to Claude on the “pause the enterprise campaign” recommendation and asked it to project the possible impacts, it returned with:

Account-Level CPA Declines 10%-15% (80%) – extremely likely, as enterprise non-branded search CPA has consistently skewed 200%-300% above SMB non-branded search. Pausing enterprise has a high probability of lowering overall CPA. This decline could be lower if enterprise terms begin to match into SMB campaigns.

Account-Level Conversion Value Declines 20%-25% (75%) – extremely likely. Enterprise non-branded search conversion values are 400%-650% higher than SMB non-branded search conversion values.

Any CEO who reads that will immediately recognize that the benefits of a lower CPA pale in comparison to losing 20-25% of the total account conversion value, and the recommendation isn’t worth it. Even better – it forces the AI to “think through” the downstream implications of its recommendations in a way that the systems don’t natively do.

Give It A Priority For Context + Assessment

In defense of Claude, it (probably) didn’t know what master it was serving when it generated this audit – does it serve the ad account or the business? Should it prioritize Google’s metrics or the brand’s performance? To you or I, the answer is obvious: prioritize the business, tell Google or Meta to pound sand. But, to the AI system, it isn’t – it simply sees two competing sets of data, then picks one (usually the one with more depth to it, which is (almost always) the platform data).

My intermediate solution has been to explicitly tell the AI the “prioritization” schedule we expect it to use – business metrics/business performance first, then ad account performance metrics, then diagnostic metrics, then vanity metrics (i.e. impression count, number of clicks, total spend).

Longer-term, I think a set of agents that review from various perspectives, with the “business performance” one being given the highest priority, is likely the play. But for now, simply stating the priority order and forcing it to show its assumptions has been effective at reducing the number of confidently-wrong recommendations.

Share The Machine On The (Human-Made) Final

After I went through the original AI audit, I uploaded our final audit to Claude and had it compare what it produced to what we created, then generate rules/assessments about where it went wrong. I was genuinely surprised that – once Claude was shown the final – it was able to quickly identify where it had gone wrong and where the evaluation was miscalibrated.

Remember To Pair Smart Machines with Smart People

Even after this experience, I genuinely believe AI powered audits can be a tremendous force for good in our industry. They can help SMBs spend smarter, identify issues faster and grow their business more reliably. Those are all wonderful things!

I know the temptation is there to let machines do everything, but this entire experience was a reminder that smart people working together with smart machines is the recipe for exceptional, efficient outcomes. Resist the temptation to automate everything, and instead focus on setting the machines up for success.

Until next week,

Cheers,

Sam

Related Insights