Nimble, a New York City startup backed by $47 million in Series B funding, launched Web Search Agents today: a domain-specialized retrieval system claiming 21% better accuracy and 51% fewer tokens compared to leading AI search alternatives. Early customer Rox, an AI-native CRM company, reported a 20x reduction in token costs after deployment. Nimble has not disclosed its benchmarking methodology or which competitors it measured against.

The core mechanism is what Nimble calls self-learning retrieval. Instead of returning broad result sets and letting a language model sort through them, the system builds a specialized retrieval model per customer domain, starting optimization on the second search with no setup required. It combines proprietary indexes, live web access, and a semantic caching layer that retains domain memory over time. CEO Uri Knorovich told VentureBeat the system reduces multi-hop reasoning steps and avoids repeatedly pushing raw pages through a language model, which is where token costs accumulate in long-running enterprise agents.

The full article is worth reading for how Nimble positions its 'Harness as a Tool' architecture against conventional search API stacks, and for the specific deployment examples across investment banking, life sciences, and supply chain workflows. The deeper question it raises: as foundation models commoditize, retrieval infrastructure is becoming the real competitive layer. Nimble is betting enterprises will need dozens of specialized search agents, not one general-purpose tool.

[READ ORIGINAL →]