Parallel packages live-web search and multi-depth research into predictable per-request APIs, but citations, completeness, latency, processor choice, and downstream content rights still need evaluation.
Direct verdict
Parallel earns a benchmark slot for research-heavy agents that need a ladder from fast retrieval to deep structured investigation. Choose the cheapest processor that meets a predeclared quality bar, verify citations at the field level, and keep latency and spend ceilings around asynchronous work.
What to verify
Assemble 150 dated questions and structured enrichment tasks across known, obscure, contradictory, and recently changed facts. Run Search, Responses, and three Task processors against a fixed human-reviewed answer set. Score field accuracy, citation entailment, source diversity, freshness, unsupported claims, completeness, p50 and p95 latency, timeout recovery, cost per accepted result, and sensitivity to query phrasing.