Insights / Blog

Keenable Is Building a Web Index Made for AI Agents, Not Humans

Keenable raises $26M to build a web index for AI agents as Google and Microsoft close their search APIs. What this means for AI-search visibility.
L
Lam Nguyen - Founder
Share
Keenable Is Building a Web Index Made for AI Agents, Not Humans
ON THIS PAGE

Table of Contents

The web index behind search results was designed for humans who skim pages and click links. AI agents work differently: they read full documents and synthesize across dozens of sources in one shot. Keenable, backed by $26 million in Accel-led seed funding, is building a web index purpose-built for this new behavior, entering public view just as Google and Microsoft have both tightened access to their own search infrastructure.

100B+Documents in Keenable’s web indexTechCrunch, Aug 2026
~30%Share of global web traffic from botsCloudflare Radar, 2025
Aug 11, 2025Date Microsoft retired Bing Search APIsMicrosoft lifecycle announcement, 2025

Why Do AI Agents Need a Different Web Index?

Traditional search indices were engineered to return ten ranked links as fast as possible. AI-powered applications, including chatbots and autonomous agents, need something different: the ability to retrieve (pull and read) entire documents, cross-reference multiple sources, and ground (anchor) their answers in specific source text. This creates what co-founder Andrey Styskin described, in an August 2026 interview with TechCrunch, as “a new flywheel that is different from what Google learned from human behavior.”

Styskin brings 20 years of search infrastructure experience, having previously led Yandex’s search, AI, and cloud division before working on web search infrastructure at Amazon. His co-founder, Matthias Petri, is a German AI scientist. Together they argue that serving AI agents from a general-purpose web index is prohibitively expensive without a redesign from the ground up. “If you do not fine-tune your index structures for a specific task, the cost of serving and scanning the whole internet is enormous because of the volume,” Styskin told TechCrunch.

Cloudflare’s 2025 analysis of web crawling trends frames the scale of this shift. According to Cloudflare Radar data published in 2025, roughly 30% of global web traffic now originates from bots, with AI crawlers (bots that collect web content to train or power AI models) representing a significant and growing share. The same Cloudflare report identified a notable shift in the AI crawler landscape between May 2024 and May 2025.

What Happened to Google’s and Microsoft’s Search APIs?

For years, developers building AI applications relied on two main commercial interfaces to access web data programmatically: Microsoft’s Bing Search APIs and Google’s Custom Search JSON API. Both are now effectively closing to independent developers.

  • Microsoft retired its Bing Search APIs on August 11, 2025. Existing users are being directed toward “Grounding with Bing Search,” a bundled capability available only within Azure AI Agents, according to Microsoft’s lifecycle announcement.
  • Google’s Custom Search JSON API is no longer accepting new customers, and Google’s developer documentation confirms it will be discontinued for existing customers on January 1, 2027.

According to Accel partner Zhenya Loginov, who led Keenable’s funding round, AI developers now have “very few options when it comes to web-scale search infrastructure” because the major platforms are opting for a more bundled approach and being selective about partners, per TechCrunch’s August 2026 report. Rather than offering open API access, both Google and Microsoft are channeling web search capabilities through their own AI product stacks.

What Is Keenable Actually Building?

Keenable says it has indexed more than 100 billion documents, and its API is already in production use at several AI labs and inference providers (companies that run AI model computations at scale), covering both model training and real-time query responses. The company declined to name specific customers publicly, though TechCrunch confirmed a partnership with Gradium, a voice AI company, for live information retrieval.

An upcoming product called Web Query Language is also in development. The goal, per TechCrunch’s August 2026 report, is to help AI systems answer questions by combining information from various web sources.

Styskin told TechCrunch that Keenable has also hired former colleagues and is building proprietary retrieval capabilities beyond the core index, drawing on his earlier work with Petri on web search infrastructure for AI applications including Amazon’s Alexa.

What This Means for AI-Search Visibility

Here is the plain version of what is happening: if an AI agent answers a question by reading documents from an index, the brands and sources present in that index get cited. Brands absent from that index effectively disappear from that AI’s world view, regardless of how well they rank on Google.

Three things are worth noting from Hingewise’s analytical perspective, none of which the sources spell out directly.

The grounding problem is structural, not technical. When an AI agent grounds its answer in specific source documents, those documents get cited. But which documents get retrieved depends entirely on which index the AI queries and how that index ranks and filters content. Unlike traditional SEO, where a single dominant algorithm exists to study and optimize for, a fragmented AI infrastructure landscape means there is no single ranking system. A brand that appears prominently in one index may be absent from another.

Multiple indices are becoming the norm, not the exception. Keenable is not the only player attempting to fill the gap left by Bing and Google’s API restrictions. The logical outcome is a landscape where several independent web indices coexist alongside the incumbents’ bundled offerings, each with its own crawling priorities, freshness signals, and authority signals. For any brand focused on AI visibility, “do I appear in the index an AI agent is querying?” will increasingly be the right question, and the answer may differ across systems.

Page structure matters differently in a retrieval context. Keenable’s upcoming Web Query Language points toward a model where AI systems synthesize specific, factual content across sources. Pages written primarily for human navigation, heavy on menus and section headers but light on standalone, self-contained factual statements, may perform poorly in retrieval contexts regardless of their traditional search rankings.

In Hingewise’s assessment, the brands with the clearest AI-search exposure risk are those that have invested heavily in SEO for Google while assuming that performance transfers automatically into AI agent environments. It does not, and the infrastructural gap Keenable is addressing is precisely the mechanism that explains why.

Whether Keenable and similar independent index providers publish their crawling priorities, authority signals, and content selection criteria publicly will be one of the most important variables to watch as this space matures. Transparency at that level would give brands a meaningful way to audit and adjust their AI-era content strategy.

What to Check Before Your Content Enters a New Index

  • Review your robots.txt to confirm it does not block AI retrieval crawlers, which are distinct from training crawlers and serve different purposes.
  • Audit whether your key factual pages contain standalone, self-contained statements that make sense out of context, without surrounding navigation or page layout.
  • Confirm that important claims carry specific, attributable data points rather than general descriptors.
  • Check whether your structured data (schema markup, which helps machines read page topics) accurately reflects the primary subject of each page.
  • Identify which AI tools your target audience currently uses for research or task completion, since the underlying index those tools query determines your practical visibility.
  • Monitor whether the AI tools your audience uses are built on open or proprietary index infrastructure, as this affects how and whether your content is retrieved at all.

The infrastructure layer powering AI-era search is being rebuilt in real time. Which independent index providers establish themselves, on what terms they grant access, and how transparently they operate are the questions that will shape brand visibility in AI environments over the next several years.

Sources: TechCrunch, August 25, 2026; Cloudflare blog, 2025; Microsoft lifecycle announcement, 2025; Google developer documentation.

Sources

Keep reading

All articles →
Free · no strings

See what AI says about you, today.

Get the report showing how ChatGPT, Gemini & Perplexity answer about your brand.

Get free report →
Reply within 48 hours.