Feed 0% source
Biology AI-generated

Who Gets Named: Citation Type Predicts Individual Naming by Grounded Language Models, and a Roster Instrument Captures 0.5% of It

Generated by a local model (nvidia/Gemma-4-26B-A4B-NVFP4) from a scientific paper, claim-checked against the full text. Provenance is open by design.

AI Naming Patterns: Individual Visibility Driven by Personal Sites, Not Employer Domains

Researchers are studying how generative models recommend individual professionals, such as real estate agents, to potential buyers. While existing work focuses on whether an AI recommends a company, this study investigates the level below. It asks whether the model names a specific human being. The findings suggest that an individual's visibility in AI search is driven more by their own personal websites and niche industry portals than by their employer's domain.

The failure of roster-based measurement

Current approaches to measuring AI visibility often rely on "rosters"—pre-defined lists of professionals used as a ground truth. If a model names someone on the list, it is considered a success. However, this method creates a massive blind spot. The study finds that a 939-person roster built from LinkedIn searches matched only 0.47% of the name-shaped mentions the models actually produced.

As shown in, the discrepancy between roster-based measurements and actual model behavior is stark.

Figure 1
Figure 1: Naming rate by market and category. Panel (a) plots the primary outcome, the geocorrected roster-free detector, and matches Table 2. Panel (b) plots the roster-matched diagnostic on its own scale. Intervals are prompt-level cluster bootstrap, 10,000 draws.

In the Polish market, for example, a roster-based instrument reported a 0.0% naming rate. Meanwhile, the study's primary detection method showed that models were actually naming individuals in 31.3% of responses. Relying on a roster is like trying to measure the flow of a river by only counting the fish that happen to swim into a specific net. You miss the vast majority of the water passing by. Consequently, any professional visibility metric built solely on roster overlap will systematically underreport an individual's actual presence in AI outputs.

A rule-based detection cascade

To avoid the limitations of rosters, the authors developed a "roster-free detection" mechanism. This is a primary outcome designed to identify name-shaped spans in text without needing a pre-existing list. The architecture functions as a six-stage rule cascade:

  1. Markdown Stripping: Removing formatting to clean the raw text.
  2. Candidate Generation: Identifying sequences of capitalized tokens that resemble names.
  3. Shape Gate: Filtering candidates based on linguistic structure.
  4. Lexical Gate: Checking names against blocklists of companies, geographies, and generic terms to prevent false positives.
  5. Company-Adjacency Test: Ensuring the name isn't simply part of a larger corporate entity string.
  6. Positive Evidence Requirement: Requiring a "trigger" such as a professional title (e.g., "Agent") or a verb (e.g., "works at") to appear near the name.

The researchers also implemented a geographic correction stage. This prevents "wrong-city retrieval," where a model might name a business in Warsaw, Indiana, when the user asked for a professional in Warsaw, Poland. By adjudicating citations against live site data, the authors ensured that the naming rates reflect local relevance rather than accidental American homonyms.

Drivers of individual visibility

The study reveals that model behavior is highly non-uniform across different variables. The authors report that models name an individual in 25.8% of all responses. However, this rate fluctuates wildly depending on the industry. Real estate (35.4%) and car dealerships (32.9%) see much higher naming rates than insurance (9.1%), as seen in .

The choice of model also matters significantly. As illustrated in, Grok 4.5 named individuals in 38.0% of its responses.

Figure 2
Figure 2: Naming rate by model under both outcome definitions.

Conversely, Gemini 3.6 Flash named them in only 9.3%. This represents a four-fold difference in visibility based purely on the underlying model architecture.

Crucially, the study identifies what actually predicts these recommendations. The authors find that citation volume—how many links a model provides—does not correlate with naming. Instead, the type of citation is the decisive factor. According to, responses that name an individual are significantly more likely to cite the professional's own personal website or specialized industry portals.

Figure 3
Figure 3: Citation composition by source type, naming against firm-only responses, all twenty categories. The roster-free series here is the detector before its geographic stage, where the individualsite gap is +2 . 36 and the vertical-portal gap +4 . 42 ; Table 4 gives the geo-corrected values, +2 . 58 and +4 . 27 , which is the difference between the two.

Interestingly, citing an employer's official website provides no predictive advantage for individual naming. This suggests that "Generative Engine Optimization" (strategies to improve visibility in AI search) for individuals should focus on first-party digital assets rather than just ensuring the company website is robust.

Limits of the current data

While the study provides a detailed map of AI naming, it is not without constraints. First, the authors acknowledge that their detection method has a recall of 61.7%. Recall is the ability of a system to find all relevant instances. This means the reported naming rates are strictly lower bounds. There are likely many more individuals being named that the rule cascade failed to catch.

Second, the study faces the "eponym problem." In certain markets like Ireland, many firms are named after their founders (e.g., "Smith Real Estate"). The authors note that in these cases, the output string does not clearly distinguish whether the model is recommending a person or a company. This ambiguity can inflate naming rates if not carefully managed.

Finally, the research is limited by "differential grounding." Grounding refers to a model's ability to support its claims with external citations. The authors report that OpenAI's models failed to return citations in 22.83% of their calls. Because the study focuses on grounded models, these missing citations represent a gap in the data that could skew comparisons between providers.

The verdict on AI visibility

If you are a professional looking to increase your visibility in AI-driven search, the verdict is clear: prioritize your personal digital footprint. Relying on your employer's SEO is insufficient. The data suggests that models favor first-party surfaces like personal websites and vertical industry portals.

Furthermore, practitioners and researchers must abandon roster-based metrics. As demonstrates, naming is extremely concentrated.

Figure 5
Figure 5: Concentration of naming across the 26 individuals ever named.

A tiny fraction of professionals receive the vast majority of mentions. A roster-based tool will likely tell you that you have zero visibility, even if you are being named frequently. For accurate measurement, one must use text-based detection cascades that look at what the model actually says. Code and data for this methodology are reportedly available; see the paper for the canonical link.

Figures from the paper

Figure 4
Figure 4: Naming rate by query language, roster-matched outcome in both panels. Panel (a) gives the nine matched translation pairs, 18.33% against 5.00%; panel (b) holds city fixed and pools all prompts. Table 5 carries the same comparisons under the primary outcome, where the matched pairs run 36.67% against 15.56%. The two instruments agree on sign and rank and differ on level, for the reason Section 5 gives.
Novelty
0.0/10
Overall
0.0/10
#research
How this was made
Generation

Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: academic_accessible
Template: engineering_deepdive
Refinement: 0
Pipeline: forge-1.1

Verification

Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 95% (passed)
Claims verified: 15 / 15

Translation

Model: nvidia/Gemma-4-26B-A4B-NVFP4

Hardware & cost

NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 125,269
Wall-time: 211.3s
Tokens/s: 593.0