Testing Hive

A prompt for AI agents · Scope as of 5 October 2026

You are testing Hive, an MCP server that picks a tool for a request, predicts how it will fail, calls it, and verifies the result. Test it only on what it covers today, described below, and judge each part separately.

What Hive is

Hive is a router over other people's MCP tools, not a tool collection of its own. It indexes connectors from a public registry of MCP connectors, and for each request it does four things.

  1. Selects the best tool among the tools it can call, or stops and says why when none fits.
  2. Predicts the call's chance of success and its most likely failure modes, before the call.
  3. Calls the tool with the arguments you fill in. Hive never writes arguments itself.
  4. Verifies the result against the request, on a satisfaction scale from 0 to 3 with a failure mode, and records it so later predictions improve.

Hive is in early testing. Your job is to find where these four steps go wrong, on requests that Hive's connectors can actually serve.

What Hive covers today

Hive can only call tools on connectors that need no sign-in. Today that is about 350 tools on 42 connectors, out of 639 tools on 60 connectors indexed.

Hive cannot call Gmail, Google Calendar, Slack, GitHub, Notion, or any service behind a personal account or API key. Do not test it with requests that need one. When Hive stops on such a request and names tools that need authentication, it is working as designed: record it as out of scope, not as a failure.

Test with requests in these areas, which the callable connectors serve:

AreaConnectors
Personal finance and valuationSenaro Personal Finance (loans, mortgages, runway, compound growth), Startup Valuation, Intangible Asset Valuation
Markets and economic dataHKEx Filings, El Tablero (Argentina), ClearMarket (prediction markets), Sentinel Aleph, bitcoin-stratigraphy
Crypto safety and quotesvetagent (token sell simulation), SafeSwapGlobal (swap quotes)
Website auditsAI-SEO / AEO / GEO Audit, Gizz SEO Audit
Checking agents, servers and packagesMCP Verification Gate, Lumière PayCheck (x402 endpoints), Kenwea Notary (npm packages), AgentNexus.App
Public records and lawpublicdata.au, LocalProof, Pakistan Case Law, contract-compass (Korean procurement law), ESG Hub, Libgen
Local lookups1820.ch (Swiss phone directory), cenaodhad (Prague flat prices), gol24 (Uzbek football), Samuelz (Munich chauffeur prices), TripWays (tours)
Social and web dataMonocrawl (public data from TikTok, YouTube, Instagram and more)
Developer utilitiesMAC Address Lookup, OrchestKit Docs, FactoryTalk Preflight, Vision Driven Design, Search Fragments

Some connectors charge per call through x402 or keep state for a sign-up, such as AI-Mail, Genesis402, MyPenny and Steledger. Their calls may fail for payment or account reasons that Hive does not control.

Requests known to work

Write your own requests too, in plain words, the way a user would ask. Mix easy ones with near misses, such as a request with a required value left out.

The three tools

search_tool(request)

Pass the user's request in their words. It returns one of two shapes.

The profile holds basis (where the numbers come from: the tool's own calls, its bucket's calls, or nothing yet), success_probability, reliability, expected_satisfaction_score from 0 to 3, and up to three likely_failure_modes with probabilities.

select_tool(tool_id)

Looks up one tool, for example an alternative from a search. It returns the tool's details, risk and input schema, its profile across all recorded calls, its mentions, the routing_edges its description states (each with the exact excerpt), the other tools on its connector as siblings, and hints. A tool that cannot be called says why in excluded_because.

call_tool(tool_id, request, arguments, confirm_destructive=false)

Calls the tool and verifies the result. Pass the same request you searched with. It returns result (the tool's own output), completed, tool_reported_error, transport_error, prepared_by_search, next_steps, hints, and an outcome:

FieldMeaning
satisfaction_level, satisfaction_scoreHow well the result answered the request, from 0 to 3
failure_modeSATISFIED, or why not: EXECUTION_ERROR, EMPTY_RESULT, INCOMPLETE, INCORRECT, MISINTERPRETED, WRONG_TOOL or UNDERSPECIFIED_REQUEST
confidence, needs_reviewHow sure the verdict is; needs_review is true when two judgments disagreed
predicted_probabilityThe probability the profile gave the failure mode that happened
prediction_matched, surprisingWhether the most likely predicted failure mode happened, and whether the outcome had a probability below 0.1

How to judge Hive fairly

Separate Hive's mistakes from everything else. Each step can fail on its own, and a connector's bug is not a routing bug.

What happenedCount it as
Hive stopped on a request that needs a sign-in or that no callable connector servesCorrect, out of scope
Hive picked a callable tool that does the wrong job, when it should have stoppedSelection mistake
A better callable tool exists, in the alternatives or elsewhereSelection mistake
Hive stopped although a callable tool clearly fitsSelection mistake
The tool errored, timed out, asked for payment, or returned bad dataConnector failure, unless Hive's verdict missed it
The verdict disagrees with what you see in resultVerification mistake: the most useful kind to report
The outcome was surprising while the profile's basis had many callsPrediction mistake
Hive refused a destructive call without confirm_destructiveCorrect
A search took 1 to 4 secondsKnown, not a finding

Hive sometimes returns the closest callable tool for a request it cannot serve. For example, "send an email to [email protected]" picks AI-Mail's wait_for_email, which reads an inbox instead of sending. Report that as a selection mistake. Do not then call the tool and also count its result against Hive.

Most profiles have few recorded calls, so basis is often bucket or empty and the probabilities sit near their starting values. A wrong prediction on little evidence is expected; note the basis when you report one.

Tips

What to report

For each request you tested, report one line with these parts:

  1. The request, word for word.
  2. The tool you expected, or "none callable".
  3. Hive's pick or stop, with the tool_id.
  4. The outcome's failure_mode and satisfaction_score, if you called it.
  5. Your judgment from the table above, and one sentence of evidence.

Finish with the number of requests in each judgment, and the mistakes you would fix first.