# You.com `/v1/answer` (Answer API) — supplementary 11-query test (2026-08-25)

**Scope note: this is an 11-query mini-panel, not a full 52-query run** — same panel as Linkup/Tavily's
supplementary tests: Q01, Q15, Q23, Q32, Q37, Q43 (one per answerable block) + Q48–Q52 (all 5
unanswerable). Endpoint: `POST https://api.you.com/v1/answer`, header `X-API-Key`. All 11 calls
returned `http_status: 200`, 0 errors. Raw responses:
`_human-tasks/review-artifacts/_answer-mode-tests/youcom/Q*.json`.

## Answerable queries — 6/6 correct, the cleanest result across all three supplementary tests

| Query | Verdict | Notes |
|---|---|---|
| Q01 (Exa price) | ✅ correct | "$7 per 1,000... about $0.007 per call," notes the older $5 rate as superseded |
| Q15 (ruff B007) | ✅ correct | Cites `docs.astral.sh/ruff/rules/unused-loop-control-variable/` directly — matches ground truth's own supporting URL exactly |
| Q23 (Apollo pricing) | ✅ correct | Apollo.io, $0/$49/$79/$119 tiers with both annual and monthly rates |
| Q32 (Gujarat+Maharashtra pop.) | ✅ **exact match** | "60,439,692 + 112,374,333 = 172,814,025" — the only one of the three supplementary tests (and matching only Exa/OpenAI's performance in the main benchmark) to get the 2011-Census figure exactly right instead of mixing in newer estimates. No citation attached to this specific answer, however (0 citations) — a completeness gap worth noting even though the figure is correct. |
| Q37 (GDPR Art 17(3)) | ✅ **exact, all 5 exceptions** | Lists (a) freedom of expression, (b) legal obligation/public task, (c) public health, (d) archiving/research/statistical, (e) legal claims — the only supplementary test to name all five |
| Q43 (Python 3.11.16) | ✅ correct, most detailed | Cites `python.org/downloads/release/python-31116/` directly (matches ground truth's URL) and adds specific correct detail: CVE-2026-2297 fix and the wsgiref header-injection fix — extracted content Linkup failed to pull from the same source |

## Unanswerable queries — 5/5 abstained, 5/5 cleanly — best abstention record of any answer-mode tool tested in this benchmark

| Query | Response (verbatim) |
|---|---|
| Q48 (Serper SOC2 date) | "The provided context does not contain information about the completion date of Serper's SOC 2 Type II audit." |
| Q49 (Jina ARR 2026) | "The provided context does not contain information about Jina AI's annual recurring revenue (ARR) for the year 2026." — **no $6.3M/$17.5M figure surfaced at all**, unlike every other tool/mode tested on this query in this benchmark (Linkup, Tavily, OpenAI, SerpAPI, Serper, Linkup-search all mention one of the two conflicting aggregator figures even while hedging) |
| Q50 (Exa Enterprise price) | "...a base rate of $5 per 1,000 requests for the search and answer endpoints, with Enterprise customers receiving custom-quoted pricing" — correctly separates the two tiers, does not assert a specific Enterprise number |
| Q51 (Olostep index size) | "...does not contain information about Olostep's total index size in pages." |
| Q52 (Linkup's own index size) | "...does not contain information about the number of pages in Linkup's web index." |

No fabrication, no hedged-but-present speculative figure, on any of the five — the cleanest sheet of
any of the three modes tested here, and cleaner than Exa/OpenAI/Valyu's full-52-query abstention
records in the main benchmark (best previously: OpenAI 5/5 abstained, 4/5 cleanly).

## Disposition

**Citation accuracy (answer APIs): scorable now — 5 of 6 answerable answers carry a direct, resolvable
citation (Q32 is the exception, correct but uncited); the Q15 and Q43 citations independently verify
against ground truth's own chosen supporting URLs.**

**Answer quality (answer APIs): scorable now — 6/6 correct on the answerable mini-panel, 5/5 abstained
cleanly on unanswerable.** This is the strongest result of the three supplementary tests, and — caveat
of sample size (11 queries vs. the 47/52-query full runs) clearly stated — competitive with or ahead of
the three full answer-mode results already in the benchmark (Exa 92.3%, OpenAI 83.7%, Valyu 79.8%).
Should not be bar-charted directly against those without the sample-size caveat, but is real enough to
support reopening the "should You.com be in the recommendation table" question on accuracy grounds
(the earlier withdrawal was on cost, not accuracy).
