In the rapidly evolving landscape of search engine optimization (SEO), the rise of Generative AI has introduced a chaotic variable that many digital marketers are struggling to master. For years, the industry relied on visibility scores and broad traffic metrics to gauge success. However, as AI Overviews (AIO) and LLM-driven search interfaces become the primary way users interact with the web, those metrics are proving insufficient. A recent deep-dive webinar hosted by Search Engine Journal featuring the experts from seoClarity—Mark Traphagen (VP of Product Marketing & Training), Mihir Naik (Senior Product Manager, AI), and Suraj Lalchandani (Sr. IT Project Manager)—sought to move the industry away from "inference-based" marketing and toward a rigorous, scientific methodology centered on causation. The Core Challenge: Correlation vs. Causation The fundamental problem in AI search today is that most teams are operating on gut feelings. A brand sees a spike in traffic and attributes it to a recent content update, only to find that the AI engine shifted its underlying model simultaneously. The seoClarity team provided a stark example of what true scientific rigor looks like: they added FAQ sections to a set of test pages and witnessed a measurable lift in AI citations. To confirm this wasn’t a mere coincidence, they removed the FAQ sections. The citations dropped immediately. This "reversion" is the gold standard of proof in A/B testing—the ability to reproduce the result by removing the variable. Without this level of testing, most teams are simply observing noise. "Visibility scores tell you if you showed up," explained the panel. "Page-level performance and split testing tell you if what you did actually mattered." Chronology of the Measurement Shift The shift toward more accurate measurement began in earnest on June 3, when Google launched dedicated Search Console reports for AI Overviews and AI Mode. This development represents a watershed moment for SEO professionals. For the first time, site owners can see—page by page—how often their URLs are surfacing within Google’s AI-generated results. Suraj Lalchandani described this as the "biggest measurement upgrade" in the short history of AI search. Previously, the industry was forced to rely on sampling and broad inferences. Now, while still limited, Google is providing direct, first-party data. However, the team cautioned that this is only one piece of the puzzle. While Google’s data is invaluable, it does not cover the broader ecosystem. ChatGPT, Claude, and Perplexity still remain "black boxes" that require sophisticated, third-party structured tracking to understand. The takeaway for practitioners is clear: use the new Search Console reports as a foundation, but do not mistake them for a complete picture of your brand’s AI search footprint. Supporting Data: Building the "Golden Set" One of the most actionable insights from the session was the strategy of building a "golden set of prompts." Rather than testing randomly, the seoClarity team advocates for building a library of prompts that mirror the full customer journey—from initial awareness to final retention. These prompts are then sorted into tiers: Tier 1: These are the "low-hanging fruit." The brand is already relevant to the topic, but the AI simply hasn’t been provided with a high-quality link worth citing. These are the easiest wins to secure. Tier 2: A heavier lift that requires more significant content or structural adjustments. Discarded Prompts: The team noted that some prompts are simply not worth the investment of time or resources, a controversial but pragmatic stance that surprised many attendees. By sequencing these tests, teams can secure early, quick wins that build the internal political capital necessary to advocate for more intensive, long-term testing strategies later. How to Run a Split Test on an LLM A recurring question in the session was how to conduct an A/B test when you cannot force an LLM to split its traffic. The solution is to move away from traffic-based splits and toward control group methodology. By identifying a set of correlated pages, marketers can create a "noise filter" that accounts for model updates and algorithmic shifts. Without a control group, the team argues that any result is effectively guesswork. Crucially, the team emphasized the importance of timing. AI search does not react in real-time. The methodology requires a strict baseline period followed by a minimum "test window." If a team cuts this window too short, they risk misinterpreting temporary fluctuations as permanent trends. The Results: Lessons from Real-World Tests The webinar shared the outcomes of three distinct client tests, providing a rare look at what actually moves the needle: The FAQ Success: As mentioned, adding FAQ sections successfully increased citations, and the removal confirmed the causality. The Meta Description Test: This test did not yield the expected results, providing a lesson on where effort should not be spent. Listicle Formatting: Like the meta description test, this failed to move the needle in the way the team hypothesized. Mihir Naik framed these "failures" as successes, arguing that knowing what doesn’t work is just as valuable as knowing what does. "Every result is a win because you have evidence instead of guesses," he noted. Q&A: Expert Insights on AI Authority and Implementation The session concluded with a Q&A that addressed the most pressing concerns in the SEO community: Defining AI Authority There is no single "authority metric" for AI. Instead, the team suggests stacking multiple signals, such as citation share on high-value prompts and cross-engine consistency. If you are consistently cited across Google, ChatGPT, and Perplexity for a specific category of questions, you have effectively established yourself as the authoritative source. The Collapsible FAQ Debate Many web developers hide content behind "read more" toggles to save space. The panel warned that implementation is everything. If the content is not technically accessible to the crawler when in a collapsed state, the AI will ignore it. The simple advice: if you aren’t sure how your specific implementation is being parsed, test it. The Role of Traditional SEO Is classic, technical SEO still relevant? The answer was a resounding "yes." Traphagen noted that clients with technically sound, well-optimized sites perform significantly better in AI search. AI optimization is not a replacement for traditional SEO; it is a specialized layer built on top of it. Implications for the Future of Search The overarching message of the webinar is that we have entered an era of "evidence-based AEO" (Answer Engine Optimization). The days of optimizing for a list of blue links are transitioning into an era where brands must optimize for their presence within generated answers. For digital marketers, the implications are profound: Data Literacy is Mandatory: You must understand how to construct control groups and interpret baseline data to separate signal from noise. Prioritization is Key: Use a tiered prompt strategy to maximize resources and secure early, demonstrable wins. Causality Over Correlation: Stop reacting to every fluctuation in your rankings. Unless you can prove a change led to a specific outcome through a controlled test, you are likely chasing ghosts. As AI engines continue to iterate, the gap between those who rely on "SEO lore" and those who rely on "experimental data" will only widen. By following the blueprint provided by the seoClarity team—focusing on golden prompt sets, rigorous control groups, and clear causality—brands can move from the uncertainty of the current AI search environment to a position of measurable, predictable authority. To view the full methodology, including the specific blueprints for schema and markdown testing, practitioners are encouraged to watch the full webinar on demand. In an industry defined by constant change, having a reliable testing framework is the only way to ensure your brand remains at the center of the AI-generated conversation. Post navigation The Blueprint for Brand Success: Navigating the Bluesky Ecosystem The Living Room Revolution: How YouTube’s Shift to TV is Rewriting the Rules of Digital Marketing