Hark, the Brett Adcock-founded AI startup backed by $700 million at a $6 billion valuation, launched Handoff today: a computer use agent that autonomously navigates the open web to book flights, order food, and message job candidates. The company claims a 97.7 score on Online-Mind2Web, beating GPT 5.4 at 92.8, Claude Opus 4.8 at 84.1, and Gemini 2.5 Pro at 69. Pricing is $0.18 per million input tokens and $2.37 per million output tokens, roughly one-tenth the cost of GPT 5.5. Per-turn latency is 0.8 seconds. Each request runs in a dedicated virtual machine with its own browser, file system, and terminal.
The benchmark story has holes. Hark's comparisons are against last-generation models. GPT-5.6, Anthropic's Opus 5, DeepSeek V4, Kimi K3, and Qwen3.8-Max are all absent. On OSWorld 2.0, Opus 5 scores 70.6% versus 55.7% for the Opus 4.8 Hark used as its baseline. On one of Hark's own three benchmarks, WebTailBench v2, GPT 5.5 outscores Handoff 72.3 to 68.6. Two of the three benchmarks were run inside Hark's own harness with Hark's own LLM judge. The base model Handoff was built on has not been disclosed. Pre-training hasn't happened yet. Enterprise security details won't arrive until end of summer.
The full article is worth reading for two reasons. First, the benchmark methodology section exposes exactly how frontier AI companies frame performance claims, and the specifics here are instructive beyond just Hark. Second, Adcock's track record adds context: he previously founded Vettery, sold for roughly $100 million in 2018, then Archer Aviation, then Figure AI, where a Fortune investigation found a BMW 'fleet' of humanoid robots amounted to a single unit. Hark's models are reportedly being trained on Figure robots. Adcock runs both companies simultaneously. The pricing advantage over current frontier models is real. Whether the performance claims hold up against 2026's strongest systems is still an open question.
[READ ORIGINAL →]