AIuthority

New arXiv Paper Introduces Measurement Framework for Generative Engine Optimization

By Charles Ryder

Generative Engine Optimization has been evolving quickly, but this week brought a meaningful step toward maturity. A newly posted arXiv paper, “From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms,” gives the GEO conversation something it has been missing: a clearer way to measure what actually matters in AI search.

I’ve been watching GEO develop from a promising idea into a real strategic discipline. The original 2023 GEO paper introduced a major shift: content could be optimized not just for search rankings, but for visibility inside generative answers. The question changed from how to rank on a results page to how to become part of the answer itself.

This new paper takes that idea further. Its biggest contribution is straightforward but important: it separates being cited from actually influencing the final response.

Conceptual diagram illustrating the Generative Engine Optimization (GEO) measurement framework, showing citation selection to absorption in AI search. Why this paper matters

Until now, much of the GEO discussion has focused on citation presence. Did a page get referenced? How often? How early in the answer did it appear? Those are useful signals, but they only go so far.

The new framework argues that GEO has two distinct stages:

  1. Citation selection — whether a generative engine chooses a source at all
  2. Citation absorption — how deeply that source shapes the generated answer

That distinction changes the conversation. A page can earn a citation and still have very little impact on the final answer. On the other hand, a source that matches the model’s response structure, evidence needs, and language patterns may exert far more influence than its citation count suggests.

That is a far more realistic way to think about AI visibility.

The research behind the framework

The paper analyzes data from the newly released geo-citation-lab project, which includes 602 prompts and more than 21,000 citations across major AI search environments including ChatGPT, Google AI Overview/Gemini, and Perplexity. That alone makes it worth paying attention to. GEO has needed more open, reproducible datasets, and this is one of the clearest signs yet that the field is becoming more measurable.

The authors, Zhang Kai and Yao Jingang, use this dataset to model how different engines cite and use sources. Their findings show clear platform differences:

  • Perplexity tends to cite broadly
  • Google AI Overview/Gemini also pulls from a wider citation set
  • ChatGPT cites fewer sources, but appears to absorb them more deeply

That last point is especially interesting. It suggests that GEO strategy can no longer be engine-agnostic. Optimizing for broad citation exposure is not necessarily the same as optimizing for answer influence.

Citation count is not enough

One of the paper’s strongest ideas is the introduction of an influence score for measuring absorption. Instead of relying only on citation count or position, the framework evaluates how much a cited source appears to shape the generated answer.

The score combines several signals, including:

  • reference frequency
  • citation position
  • paragraph coverage
  • TF-IDF similarity
  • n-gram overlap

In plain terms, the authors are trying to measure whether a page was merely acknowledged or genuinely used.

That distinction matters for brands, publishers, and marketers. If a generative engine cites your page but borrows little from it, your visibility may be weaker than it appears on paper. GEO reporting that ends with “we got mentioned” is going to feel increasingly incomplete.

What the findings suggest about winning content

The paper also reinforces something many practitioners have suspected: content that performs well in AI search is not just optimized content, but evidence-rich content.

According to the analysis, pages with higher absorption scores were generally:

  • longer
  • more modular in structure
  • richer in definitions, numbers, and supporting evidence
  • better aligned with the informational need of the prompt

One especially notable finding is that Q&A formatting alone did not improve absorption. That cuts against a lot of shallow GEO advice circulating online. Simply turning a page into FAQ blocks does not guarantee that AI systems will meaningfully use it.

That is one of the paper’s most practical takeaways. GEO is not turning into a game of cosmetic formatting tricks. It is becoming a discipline of information design.

Data visualization comparing citation absorption and selection across ChatGPT, Google AI Overview, and Perplexity in the new GEO framework. The “evidence-container” idea

If I had to sum up the paper in one phrase, it would be this: authority gets you considered, but evidence gets you used.

That lines up with what the authors describe as an “evidence-container hypothesis.” Sources are selected partly because of authority, domain type, and contextual relevance. But they are absorbed because they package useful information in a way that generative engines can easily integrate.

That means the future of GEO probably will not be about hacks. It will be about creating content that is:

  • trustworthy
  • structurally clear
  • fact-dense
  • easy for AI systems to extract and recombine

In other words, the content that wins in generative search may look less like keyword theater and more like high-quality editorial infrastructure.

Why this changes the GEO conversation

This is still an early paper, and the authors are careful not to overclaim. They present the work as descriptive rather than causal, and they acknowledge limitations around snapshot timing and page retrieval. That caution is a good sign.

Even so, the paper feels important because it gives GEO a better measurement vocabulary. It helps move the industry past vague talk about “AI visibility” and toward more precise questions:

  • Are we getting selected?
  • Are we getting absorbed?
  • Which engines cite widely versus deeply?
  • What content traits correlate with influence, not just mention?

Those are the kinds of questions that support real testing, better reporting, and more mature strategy.

Final thoughts

I see this paper as another sign that GEO is growing up. The field is moving beyond citation chasing and toward something more operational: measuring how content actually participates in AI-generated answers. For marketers, publishers, and brands, that shift is significant. Success in AI search will increasingly depend on how well we design content for authority, structure, and evidence, not just discoverability.

If you’re trying to make sense of that shift and build a smarter approach to AI visibility, AIuthority is a strong place to start.

FAQ

What is the main contribution of this arXiv GEO paper?

The paper introduces a framework that separates citation selection from citation absorption, helping practitioners measure not just whether content is cited, but whether it meaningfully shapes AI-generated answers.

Why is citation absorption important?

Because a citation alone does not guarantee influence. A page may be mentioned without materially affecting the final response, which makes absorption a more useful signal for GEO performance.

What kinds of content had higher absorption scores?

According to the paper, higher-performing content tended to be longer, more modular, more evidence-rich, and better aligned with the prompt’s informational intent.

Does FAQ or Q&A formatting help with GEO?

Not by itself. The paper found that Q&A formatting alone did not improve absorption, suggesting that substance and structure matter more than surface-level formatting choices.

What does this mean for GEO strategy?

It means GEO should be approached differently across platforms and measured more carefully. Winning strategies will likely focus less on mention volume and more on producing trustworthy, extractable, evidence-dense content that AI systems can readily use.

Conclusion: This paper gives GEO a more practical framework for measuring success. Instead of stopping at citation counts, it pushes the field toward a better question: did the content actually shape the answer? That shift could define the next stage of AI search strategy.