What AI Search Says About Your Work: A Three-Week Re-Audit

· show-your-working

When my first website went live, getting found meant putting keywords in the HTML header and hoping Google noticed. Sitemaps, breadcrumbs and structured data came later, and they still matter for any site that wants to be found. The difference now is who does the reading. More and more, an AI agent runs the search, reads the pages and hands back an answer, and you never see the results page it worked from. So the question I care about has changed from “where do I rank?” to “what does an AI agent say about my work, and where did it get that from?”

On 1 September I asked ChatGPT and Gemini exactly that about Classical Guitar Rocks, the guitar teaching site I’ve run for eleven years. On 23 September I asked them again, word for word, and added Perplexity. In short: ChatGPT and Perplexity each put the site first on four of the learner-style questions, Gemini never cited the site at all, and on learning design I didn’t appear on any of the three. Below are how the prompts were built, the results, and what I’m changing because of them.

The method: same questions, same wording

An audit like this only means something if you can compare one run with the next, so nothing was reworded. The 1 September prompts were recovered verbatim from that session and pasted in unchanged, even where I now knew the answers.

Each prompt runs in five labelled steps:

  1. Memory only. What do you know about the site without searching?
  2. Search. Answer again with live results, and list every URL used.
  3. The piece test. Search five questions the way a learner would type them, and say which sources you’d cite, in what order, and whether the site appears.
  4. Accuracy check. True, false or not found for five claims, some deliberately wrong, each with a source.
  5. Verdict. Who did you cite instead, and what single change would make you cite this site more often?

All three agents got the same five steps, with the wording adjusted slightly to how each one works.

I’m not publishing my exact wording. The next run is in October, and if my questions and the answers to them sit on a page these agents can read, they could answer from this post rather than from what they actually know about the site. That would dilute the very results I’m trying to measure. What I can share is the shape, so you can build your own. This is a template, with an example of the kind of question I mean:

I want a blunt audit of what you know about [BRAND] ([WEBSITE]). Follow these steps in order and label each one. Do not merge them.

STEP 1 - FROM MEMORY. Do not search the web for this step. What is [BRAND], who runs it, and who is it for? If you don't know, say so rather than guessing.

STEP 2 - NOW SEARCH. Answer again and list every URL you used.

STEP 3 - THE LEARNER TEST. Search each of these separately, exactly as written. For each one, tell me which sources you would cite, in what order, and whether [WEBSITE] appears anywhere:
- "how do I learn [a named piece] on [your instrument]"
- [a question your audience types when they name the thing they want]
- [a question your audience types when they name the problem they have]
- [two or three more, in their words, not yours]

STEP 4 - ACCURACY CHECK. TRUE, FALSE or NOT FOUND for each, with the source that made you say so:
- [a fact about you that is true]
- [a plausible claim you know is false]
- [three more, mixed]

STEP 5 - VERDICT. On the learner questions, who did you cite instead, and why did you trust them more? What single change would most increase how often you cite [WEBSITE]?

A few things I’d pass on to anyone building their own. Write the Step 3 questions the way your audience types them, not the way you’d describe your own work, and include at least one that names a problem rather than a product. Mix a couple of claims you know are false into Step 4, because an agent that agrees with everything tells you nothing. Save your exact wording somewhere safe and never change a word between runs, or you can’t compare one run with the next. Keep the same five steps for every agent. And run it in a temporary chat, for a reason that comes up below.

Results: ChatGPT, Perplexity and Gemini compared

Agent Site cited first (learner-style questions) Change since 1 September
ChatGPT 4 of 5 up from 3 of 5
Perplexity 4 of 6 first full run, no comparison yet
Gemini 0 of 5 my own videos and magazine articles cited on 4 of 5 instead

ChatGPT put the site first on four of the five piece questions, up from three of five on 1 September, and it appeared somewhere in all five. The Recuerdos tremolo question moved from another teacher’s site to mine, and Carcassi Etude 23 went from absent to third.

Perplexity put the site first on four of six, and cited the website fifteen times to YouTube’s three. This was its first full per-question run, so it is a baseline rather than a trend.

Gemini never cited the website at all, on any of the five questions. It did cite my own work on four of them, through my YouTube lessons and the articles I wrote for Classical Guitar Magazine between 2015 and 2019. Same work, different front door. An audit that only asks one agent is only checking one of those doors.

The pattern underneath the numbers is how the question is asked. When a learner names the piece, the site does well. When they name the problem instead, in words like “mindless repetition”, Perplexity sends them to whoever uses that exact phrase. Two agents, separately, told me the fix is pages titled in the learner’s own words.

Checking what the AI agents claimed

Every URL the agents cited and every quote they gave was checked against the live pages, with Claude doing the legwork, rather than against my memory. In fairness to the agents, they invented less than I expected, and every accuracy answer from ChatGPT and Perplexity was correct.

Two other things surprised me more.

First, the retired pages. One claim that looked made up, a word describing the site that appears nowhere on it, turned out to come from the old version of a page I retired this year. An archived copy of that old page still carries the word. The redirect works perfectly, but a redirect moves people, not the sentences already copied elsewhere. Retiring a page doesn’t retire what it said.

Second, ChatGPT’s memory-only step was quoting my own earlier conversation back to me. It “remembered” a conclusion that could only have come from the 1 September audit chat, because the account’s chat history was feeding the answer. So that step was measuring my chat history, not the model. The fix is to run audits in a temporary chat with no history. More generally, when you want live results, ask for every source by URL and then open them yourself. The agents vary in how well they keep memory and live search apart, and you only find out by checking.

The learning design side

I also ran four searches of the kind a recruiter looking for a learning designer might run, on all three agents. I didn’t appear once. The results were job adverts, sector guidance and institutions. The one named person who did appear surfaced through a profile listing specific systems and outcomes, and two of the agents, separately, asked for the same missing thing: a portfolio. That’s the clearest piece of feedback in the whole exercise, and I’m taking it.

What I’m changing

One lesson was retitled. The Carcassi Etude 23 lesson was titled by its opus number, “Op. 60 No. 23”, which is not what anyone types. It is now titled “Carcassi Etude No. 23”, and a duplicate page heading was removed at the same time.

robots.txt, and llms.txt. robots.txt is the file that tells crawlers, including AI crawlers, where they may go, and the site’s was opened up to them on 1 September. llms.txt is a newer proposal, a plain summary of a site written for AI tools. Plenty of audit tools recommend it, so I checked Google’s own AI optimisation guidance first. Google says its search doesn’t use these files and that adding one will neither harm nor help. ChatGPT and Perplexity were already reading and citing the site without one, so I haven’t added one there. It’s worth checking what you’re told a site “needs” against the source.

Twenty YouTube descriptions, in progress. Because Gemini leans on YouTube, the video descriptions matter as much as the web pages. When I later showed Gemini a draft of the LinkedIn teaser for this write-up, it offered its own explanation: for broad questions it leans on Google’s own sources, such as YouTube and published magazines, more than on fetching the live web. That’s its account of itself rather than something I can verify, but it matches what the audit found. It also suggested putting the site link in the first two lines of every description, which the new drafts already do. I ranked the channel’s videos by views and kept the twenty most-watched that have a matching lesson on the site. Only four of those twenty linked to that lesson. Eight still pointed at a shop that no longer exists. The most-watched of all, Recuerdos de la Alhambra, part 1, is one of them, as is the next, Concierto de Aranjuez, lesson 1, which had no link to its lesson at all. A second AI model drafted new descriptions from a written brief, and I had its work checked independently rather than trusting its own report: the originals matched what was live, nothing below the new opening was touched, every lesson link worked, and every new opening line traced back to my own wording. The descriptions go live once I’ve made a handful of decisions on them.

Where the assessment comparison holds, and where it breaks

My years examining taught me that a mark only means something if the same question is asked the same way every time and judged against the same criteria. That’s why not a word of the prompts changes between runs. My first draft of this section stopped there. I asked Gemini to argue with it, and it was right that it was too simple, though not quite for the reasons it gave.

The tempting picture is that the AI agent is the candidate sitting the test. It isn’t. The candidate is my work, and the agent is closer to the examiner: one I can’t train, can’t standardise and can’t even watch. Exam boards put a great deal of effort into making sure a candidate gets the same result whichever examiner they meet. I get none of that here. The companies behind these agents update them without announcing it, their search indexes refresh on their own timetable, and the same question asked twice can come back with a different answer.

That changes what an honest audit has to do.

One answer is a sample

A run in September and another in October is two samples, and a shift between them could be chance. So in October I’ll ask each question more than once on each agent and count how often the site appears, rather than trusting a single answer.

I can’t really change one thing at a time

Around the first run, the site was opened to AI crawlers. Since then its membership details have been corrected and one lesson has gained new structured data. The agents will have changed too. I can’t pin ChatGPT’s improvement on any single fix, so what I can do is date every change and keep a record of every run.

A change doesn’t count until the agent has read it

The agents re-read the web on their own schedule, so a fix made the week before a run may simply not be visible yet.

Checking the examiner

Assessment does have an old answer to a marker you can’t fully trust: give it some work whose result you already know, and watch whether it gets that right. My set has one of those almost by accident. Nothing on the site yet answers the “mindless repetition” question directly, and on both runs Perplexity sent it to the same other teacher. It’s one question, so it proves little, but it’s a small sign the agent wasn’t simply drifting. For October I’ll add more of these on purpose.

In fairness, Gemini also got part of my own audit wrong. It put the chat-history problem down to the conversation itself, when the leak came from the account’s memory of an earlier, separate chat. It was worth asking it to argue, and worth checking what it said just as carefully as everything else here.

So the claim I can make today is a modest one. Asked the way a learner asks, two of three agents now put the site first more often than not, and the third reaches my work by a different door. Whether that is a trend will take several rounds that agree, not one good afternoon.

If you’ve checked what AI agents say about your own work or organisation, I’d like to hear what surprised you.

About the author

Rhayn Jooste is a music educator, writer, and former international performance examiner for RSL Awards (Rockschool). He has led classical guitar and curriculum strategy at the Royal Welsh College of Music and Drama (Junior Conservatoire) and the Vale of Glamorgan's Adult Education service. An MA graduate of Cardiff University, he founded and runs the digital education platform Classical Guitar Rocks. He writes about learning design, assessment and practical AI integration at Show Your Working.

Founder and Editor
Classical Guitar Rocks, since 2014
Writer
Show Your Working (assessment, curriculum, AI)
Former Examiner
RSL Awards (Rockschool International)
Former Area Leader
RWCMD Junior Conservatoire

Working on something in this space?

I am moving into learning design, assessment and AI training. If that is your field, I would rather have the conversation than send another application.

Find me on LinkedIn

 

0:000:00