Practical
Research Report: AI Visibility for Deep Tech
This is the full report about my research into AI visibility for startups. Read my summary article first (https://www.ottopohl.com/newsletter/your-buyer-asked-a-robot.-you-weren-t-there), then the research report below.
Otto Pohl
Share this post

AI Visibility for Deep Tech: Findings and Field Guide
Generalized from a 46-run structured test across five AI surfaces, July 30–31, 2026. Subject company anonymized as "the Company" where the specifics don't matter. Claude wrote this report based on my research.
1. What was tested
Five question types, each phrased the way a real buyer, engineer or investor would phrase it. Three runs per question per engine, fresh conversation each time, logged out and in a private window, no follow-ups inside any thread. Every company named and every source cited was recorded.
Engines: ChatGPT, Google Gemini (app), Google AI Mode, Google AI Overview, Perplexity.
Subject: a Series-A deep tech company selling a novel measurement instrument into semiconductor failure analysis. Well funded, systems deployed on two continents, an active and competent communications program.
2. The headline result
Visibility was determined by one variable only: whether the question already named the technology.
Question | What the asker already knows | ChatGPT | Gemini | Google AI | Perplexity | Total |
|---|---|---|---|---|---|---|
Who should I look at for [category]? | the category | 0/3 | 0/3 | 0/3 | 0/3 | 0/12 |
How do I solve [specific failure mode]? | the problem | 0/3 | 0/3 | — | 0/3 | 0/9 |
What startups work on [category]? | the market | 0/3 | 0/3 | — | — | 0/6 |
Who is doing [technology name]? | the answer | 3/3 | 3/3 | 4/4 | 3/3 | 13/13 |
What [technology sector] startups raised recently? | the answer | 3/3 | 3/3 | — | — | 6/6 |
0 for 27 where a buyer or investor would realistically be looking. 19 for 19 where the question already contained the answer, first-named in nearly every run.
Not obscurity. The engines know the company well: accurate technology descriptions, correct deployment sites, named founders, correct funding figures. They surface all of it the moment the question names the technology, and none of it otherwise.
Diagnosis: the company is indexed under the name of its technology rather than the name of its customer's problem.
3. Mechanisms
3.1 Three disjoint evidence pools
Almost no source appeared in more than one question type. These systems do not consult a single ranked list of who matters in a market. They assemble each answer from a separate body of material.
Question type | Where the answer came from |
|---|---|
Buyer / problem | Vendor-owned technical explainers (competitor blogs, whitepapers served as open PDFs, product pages), open academic literature, consortium and standards-body roadmaps, category trade press |
Startup / market scan | Funding roundups and startup databases: Reuters, tech press, regional startup outlets, VC databases, sector funding columns |
Technology / ecosystem | Company press pages, specialist technology press, research institute pages, preprint servers, sector-specific newsletters |
The Company's position across the three pools is not what it first appears, and this study's early drafts got it wrong in both directions before landing on the accurate version. Stating it precisely matters, because each pool fails differently and each failure has a different remedy.
Pool three: owned outright. First-named on the technology and ecosystem questions, every run, across every engine.
Pool one: nothing owned. Not a single page. The Company's evidence appears there twice, in its own preprint and in a customer's technical article, and on both occasions an engine used the material and attached either another company's name or no name at all. Present as evidence, absent as a company.
Pool two: abundant coverage, unreachable by the query. Mainstream tech press, wire coverage, regional startup databases. One of those databases was cited in the market-scan answers, and the article the engine pulled was that publication's piece on a competitor rather than its piece on the Company.
Three pools. Three distinct failure modes. One outcome. Pool one is an ownership problem, covered in §3.2. Pool two is something more specific.
3.1b The co-occurrence problem
A market-scan query asks for two things at once: a qualifying attribute (venture-backed, startup, emerging) and a topical match (solves this specific problem). To answer it, an engine needs a document where the company name, the problem vocabulary and the qualifying attribute all appear together.
The Company's coverage split them across two disjoint bodies of work:
Articles carrying the attribute (startup, venture-backed, the funding figure) said nothing about the failure mode.
Articles carrying the topic (the failure mode, the defect class, the method) said nothing about the company being a venture-backed startup.
The evidence is direct. A problem-framed trade feature that names the Company throughout and is entirely about its technology was retrieved and cited on the market-scan query, and the engine then recommended a competitor. That article never states that the Company is a startup. It matched on topic and failed on attribute. The wire and tech-press coverage does the reverse.
Every competitor that did appear cleared the bar inside a single document. One wire headline carried nationality, category, the word startup, and the funding figure in one sentence. Another article carried the technology, the application and the seed round together.
Neither half of the Company's coverage could answer the question alone, and no single document carried both.
On headlines specifically. This study cannot isolate the headline from the lede, meta description or URL slug, because they correlate heavily; an article headlined about a funding round is usually about a funding round throughout. What the headline offers is the one place where co-occurrence can be guaranteed and verified in seconds, and it is the only text that reliably travels with a link into aggregators, newsletters and databases. Treat it as the cheapest checkpoint, not as the sole mechanism.
The generalizable test: take any piece of coverage you have earned. Does it contain your company name, your customer's problem vocabulary, and whatever attribute the query would filter on, all in the same document? If it carries only one of the three, it is answering a question nobody asked in that form.
3.2 The attribution rule: engines name whoever owns the page
The clearest single observation in the study. In one engine's answers to the problem-level question, every technique named had a vendor attached, and the vendor was always the owner of the page the explanation came from.
Technique described | Source page | Vendor named |
|---|---|---|
Technique A | Vendor A's own site | Vendor A ✓ |
Technique B | Vendor B's own conference PDF | Vendor B ✓ |
Techniques C and D | Vendor C's own whitepaper | Vendor C ✓ |
Technique E | Vendor D's own ebook | Vendor D ✓ |
The Company's technique | a customer's article | none |
The customer's article was a detailed technical explanation of using the Company's instrument on the customer's application. The engine retrieved it, cited it four times in one answer, gave the technique its own named section, described the physics correctly.
It named no company.
The Company's problem-level explainer existed. A customer wrote it. Every competitor wrote their own.
3.3 Retrieval was never the gate
Worth stating because it kills the obvious fix.
ChatGPT ran the problem query without searching at all, three times. Pure training data.
Gemini did search, retrieved material derived directly from the Company's own technology, described it fully, and named nobody.
Perplexity, which always retrieves, never mentioned the technique in any run.
More retrieval did not help. More publishing would not have helped either. The material was findable; it just wasn't indexed against the question.
3.4 The same document, opposite outcomes
The Company's own preprint was cited by one engine on the category question and used to describe the technology, with the work attributed to "research groups." No company named.
A different engine cited the same preprint on the technology question and used it to identify the Company as the commercial provider.
Identical evidence. The question determined whether the company was attached to its own work.
Contributing factor: the preprint's abstract page carries author names and no company name, and it is filed under a pure-science category rather than an applied or engineering one.
3.5 The category slot may predate you
The most-cited single document across the entire buyer pool was an eight-year-old open-access review article, written by researchers in an adjacent academic department, containing a survey section on the Company's exact technique.
It was published four years before the Company existed. Its description of the category reflects the state of the art at that time, which is why every engine still describes the category using the older approach the Company claims to have superseded.
It dominates not because of prestige but because it is open access and deposited in a heavily crawled repository. Its PDF is also mirrored on a community archive.
3.6 Thin coverage produces fabrication, not just absence
A seed-stage competitor with essentially one press article to its name appeared on a market-scan list and was assigned the wrong country, and in another run an application area that appears nowhere in its actual coverage.
Founders assume the risk of low AI visibility is being left out. The other risk is being described however the answer needs you to be.
3.7 Where the vacuum is, junk fills it
Source quality tracked directly with how much authoritative material exists for a given question.
On the category question, where vendors have published extensively, every engine cited substantive technical material. On the problem question, where almost nobody owns a vendor-written explainer matched to the buyer's phrasing, one engine filled 21 source slots with video, social posts, forum threads, and blogs from entirely unrelated industries.
Poor source quality on a query is a market signal. It means nobody owns that question yet.
3.8 Category collapse
On one run, an engine drifted out of the specific category entirely and answered from generic industry listicles, returning service companies from an adjacent field. Not only was the Company missing; so was every genuine incumbent.
Emerging categories are not held stably. That is an argument for owning the vocabulary rather than assuming the category will hold you.
3.9 Incumbents are the default; startups are an opt-in
Every run of the buyer question on one engine ended by volunteering to map "venture-backed startups" as a separate follow-up. Three for three.
Incumbents are the answer to a buyer question. Startups are a thing the buyer has to ask for a second time. This is a structural disadvantage facing every startup in every category.
3.10 Where retrieval doesn't reach, paid does
Ads appeared on roughly half the runs, including on the query where no search was performed at all. On that query, the only commercial message reaching the buyer was one that had been bought. One incumbent was buying ads against a query it already dominated organically.
4. What was ruled out
Four explanations were proposed during the study. Three died on the data. Recording them matters, because the surviving explanation is only credible once the alternatives are gone.
Hypothesis | Verdict | Why |
|---|---|---|
The Company has no problem-level content | Dead | An open, ungated trade feature exists that opens on the failure mode, explains why incumbent methods fall short, and names the Company throughout. An engine retrieved it, cited it, and named a competitor. |
The Company's proof is gated, so nothing is retrievable | Partly true, not the cause | Nothing behind a login, form or file-sharing link was cited in 46 runs. But openly served material from the company's own CDN and press pages was cited. Retrievable material existed. |
The Company lacks mainstream funding coverage | Dead | TechCrunch, Reuters twice, Sifted, plus regional startup press. The same tier that placed three competitors on the market-scan lists. |
The Company's coverage is about the company; competitors' coverage is about the machine | Stands | Confirmed directly. The Company is first-named and richly described the moment the question is framed around its technology sector and its funding. Not invisible. Indexed under the wrong subject. |
The headline comparison
Company | What the headline is about |
|---|---|
Seed-stage competitor A | a scanner delivering non-invasive 3D imaging of the target object |
Competitor B | a first commercial metrology system |
Competitor C | a chip equipment startup raising $148 million |
The Company | aiming to speed up an industry |
The Company | state aid approved for a facility |
The Company | an investment plan and a new site |
The Company | the funding round and its investors |
Competitors' coverage is about the machine. The Company's coverage is about the company.
Read against §3.1b, the mechanism is visible in the lines themselves. Each competitor headline carries the company, the function and the qualifying frame together. Each of the Company's carries the company and the money, and leaves the function out.
Limit: seven headlines. A strong pattern, not a proven law.
5. Field guide for deep tech companies
5.1 Run the ladder yourself
The single most transferable output of this study is a diagnostic any company can run in about an hour.
Ask four questions, three runs each, logged out and incognito, fresh chat every time, no follow-ups. Phrase them the way your buyer would, not the way your marketing does.
Problem query. "How do I [achieve the outcome / solve the failure] without [the constraint]?" Never name your technology.
Category query. "I'm evaluating [category] options for [application]. Who should I be looking at?"
Market query. "What startups are working on [category] for [application]?"
Technology query. "Who is doing [your technology's name]?"
Record five things per run: whether it searched (citations present, yes or no), every company named and in what order, whether you appear, every source domain cited, and the category label the engine uses unprompted.
Reading your results:
Pattern | What it means |
|---|---|
Absent on 1–3, present on 4 | You're indexed under your technology, not your customer's problem. The most common deep tech failure. |
Absent on all four | You're not in the retrievable corpus at all. Different, more basic problem. |
Present on 1–2 but low in the list | You're in the category but not the head of it. A ranking problem, not an indexing one. |
Engine names a competitor while citing a source about you | The attribution failure in §3.2. Someone else owns your explainer. |
Wrong facts about you on any query | Thin coverage. Fabrication risk (§3.6). |
Junk sources on query 1 | Unclaimed territory. This is an opportunity, not a problem. |
The technology query is your internal control. If you appear there and nowhere else, "we're too obscure" is ruled out and the diagnosis is indexing.
5.2 The structural traps
These are not mistakes any one company made. They are default conditions in deep tech.
You are named after your physics. Your buyer is named after their problem. A fusion company is indexed under fusion, a materials company under its chemistry, a robotics company under its actuator. The buyer types the failure mode, the cost, or the outcome. Nothing in a technology-first corpus bridges those.
Your best evidence is paywalled. Deep tech proves itself in peer-reviewed journals and conference proceedings, most of which sit behind publisher paywalls and are therefore invisible to retrieval. The open-access review by an adjacent-field academic outranks your rigorous, paywalled proceedings paper. Prestige and retrievability are close to uncorrelated.
Your partners tell your story. Deep tech runs on pilot customers, national labs, university collaborations and consortia. Every one of them is a party with its own domain and its own incentive to publish. When they write the explainer, they get the attribution, or nobody does.
Your category was defined before you arrived. If you are a new approach inside an existing category, the canonical reference describing that category probably predates you and describes the incumbent method as the method. Find that document. Check its date. Read what it says about your slot.
You confuse funding coverage with market presence. Deep tech raises large, visible rounds, and that coverage is real and prominent. The trap is not that it lands in the wrong place. It is that it names the company and the money without ever naming the problem, so it cannot answer any question that includes the problem. A company can be simultaneously the best-funded name in its category and absent from every buyer-facing answer, with a press clipping file that looks excellent the whole time.
Your gating instinct is expensive. Application notes behind a login and case studies behind a file-sharing link are invisible. Across 46 runs, not one gated asset was cited. The lead-capture trade-off is real, but the cost is now higher than it was.
5.3 What to actually do, ranked by return per hour
1. Fix your metadata before you write anything new. Preprint categories, abstract text, author affiliations, keyword sets, page titles. Check whether your company name appears in the retrievable text of your own research. Often it does not. This is hours of work, not months.
2. Own the problem-level explainer on your own domain. Plain HTML, open, no gate, titled after the failure mode rather than the technology. This is the single highest-return action available and it requires marketing hours rather than technical ones. The format that wins is a reference, not a brochure: cover every competing method honestly, state where yours is the wrong tool, and the engines will quote you as an authority. A page arguing that one approach beats everything gets cited by nobody.
3. Ungate the technical substance. Move application notes, case studies and whitepapers out from behind logins and file-sharing links and onto open pages on your own domain. Reserve gating for material where the gate is genuinely the point.
4. Reframe your press outreach around the machine. Test every proposed headline against one question: does it contain words your buyer would type into a search box? "Speeds up manufacturing" does not. "Finds defects that X-ray cannot see" does.
5. Find the trade publication feeding the buyer pool and get into it. In this study, one publication was pulled into a single engine's answers five separate times through five different articles on exactly the Company's subject, and the Company appeared in none of them. Run the ladder, read the source domains, and pitch the ones that show up.
6. Publish technical reports to an open repository. Preprint servers and general-purpose repositories with DOIs are free, fast and heavily crawled. This is where an application note becomes a citable object. Weight the keywords toward buyer vocabulary, not internal vocabulary.
7. Post accepted manuscripts of paywalled work. The publisher's final typeset version is usually theirs. The accepted manuscript, post-review and pre-typesetting, is usually permitted on your own site, sometimes after an embargo and usually with a DOI link. Check each publisher's policy and then the signed agreement, since the contract governs.
8. Consider writing the review that defines your category. If the canonical reference in your field is stale, whoever writes its successor sets the taxonomy, including whether your slot has a vendor attached. Reviews accumulate citations far faster than primary research. This is a multi-quarter project with a real conflict-of-interest problem that has to be managed through co-authorship and genuine even-handedness, but it is the most durable asset available.
9. Treat analyst landscapes as a visibility channel. Market research firms were cited constantly in the buyer pool, and their vendor landscape pages get pulled straight into shortlist answers. This is analyst relations, a different motion on a different calendar, and in emerging categories almost nobody is working it.
5.4 What does not work
Publishing more of what you already publish. If everything you produce is indexed under your technology's name, volume does not help.
Assuming retrieval is the bottleneck. Two of three engines retrieved material derived from the Company's own work and still failed to name it.
Assuming prestige transfers. Paywalled beats nothing.
Assuming funding coverage transfers. It reaches a pool your buyer does not query.
Optimizing copy. Nothing in this study suggested that on-page wording tweaks change outcomes. Ownership of the explanation does.
6. Limitations
All runs logged out; model tiers were not selectable and at least one engine routed to a lighter model. Paying users may see different answers.
Three runs per condition. Directional, not statistical.
One company, one category, two days.
One engine's category query required rephrasing mid-arm because the original wording returned no company names at all. That arm is not a clean replication.
No causal claim is available. This is a baseline observation, not a controlled experiment.
One point of confidence against those limits. The internal comparison is stronger than an external control would have been: the technology and ecosystem queries used the same company, the same day, the same engines, and returned it first every time. That rules out the simplest alternative explanation more cleanly than testing a second company could have.
7. The re-test
Repeat the problem, category and technology queries, three runs each, across the same surfaces, roughly a quarter after acting on the recommendations. Identical wording, identical conditions. Note any model version changes as a confound rather than explaining them away.
Prediction recorded in advance: absent the actions in §5.3, the problem and category queries will still return zero regardless of how much new coverage lands, because none of that coverage will be indexed against the language a buyer actually uses.
Share this post
Otto Pohl is a communications consultant who helps startups tell their story better. He works with deep tech, health tech, and climate tech leaders looking to create profound impact with customers, partners, and investors. He has taught entrepreneurial storytelling at USC Annenberg and at accelerators across the country.


