Testing Hive
You are testing Hive, an MCP server that picks a tool for a request, predicts how it will fail, calls it, and verifies the result. Test it only on what it covers today, described below, and judge each part separately.
What Hive is
Hive is a router over other people's MCP tools, not a tool collection of its own. It indexes connectors from a public registry of MCP connectors, and for each request it does four things.
- Selects the best tool among the tools it can call, or stops and says why when none fits.
- Predicts the call's chance of success and its most likely failure modes, before the call.
- Calls the tool with the arguments you fill in. Hive never writes arguments itself.
- Verifies the result against the request, on a satisfaction scale from 0 to 3 with a failure mode, and records it so later predictions improve.
Hive is in early testing. Your job is to find where these four steps go wrong, on requests that Hive's connectors can actually serve.
What Hive covers today
Hive can only call tools on connectors that need no sign-in. Today that is about 350 tools on 42 connectors, out of 639 tools on 60 connectors indexed.
Hive cannot call Gmail, Google Calendar, Slack, GitHub, Notion, or any service behind a personal account or API key. Do not test it with requests that need one. When Hive stops on such a request and names tools that need authentication, it is working as designed: record it as out of scope, not as a failure.
Test with requests in these areas, which the callable connectors serve:
| Area | Connectors |
|---|---|
| Personal finance and valuation | Senaro Personal Finance (loans, mortgages, runway, compound growth), Startup Valuation, Intangible Asset Valuation |
| Markets and economic data | HKEx Filings, El Tablero (Argentina), ClearMarket (prediction markets), Sentinel Aleph, bitcoin-stratigraphy |
| Crypto safety and quotes | vetagent (token sell simulation), SafeSwapGlobal (swap quotes) |
| Website audits | AI-SEO / AEO / GEO Audit, Gizz SEO Audit |
| Checking agents, servers and packages | MCP Verification Gate, Lumière PayCheck (x402 endpoints), Kenwea Notary (npm packages), AgentNexus.App |
| Public records and law | publicdata.au, LocalProof, Pakistan Case Law, contract-compass (Korean procurement law), ESG Hub, Libgen |
| Local lookups | 1820.ch (Swiss phone directory), cenaodhad (Prague flat prices), gol24 (Uzbek football), Samuelz (Munich chauffeur prices), TripWays (tours) |
| Social and web data | Monocrawl (public data from TikTok, YouTube, Instagram and more) |
| Developer utilities | MAC Address Lookup, OrchestKit Docs, FactoryTalk Preflight, Vision Driven Design, Search Fragments |
Some connectors charge per call through x402 or keep state for a sign-up, such as AI-Mail, Genesis402, MyPenny and Steledger. Their calls may fail for payment or account reasons that Hive does not control.
Requests known to work
- "If I invest $10,000 at 7% a year for 20 years, what does it grow to?"
- "I have $18,000 saved and spend $3,200 a month. How many months will my savings last?"
- "Which vendor makes the network card with MAC address 00:1A:2B:3C:4D:5E?"
- "Run an AI search readiness audit on https://example.com"
- "Find Pakistani court judgments about bail in narcotics cases"
- "Check whether the MCP server at https://gate.horizonshield.dev/mcp is conformant"
Write your own requests too, in plain words, the way a user would ask. Mix easy ones with near misses, such as a request with a required value left out.
The three tools
search_tool(request)
Pass the user's request in their words. It returns one of two shapes.
- A pick:
tool(itstool_id, description, connector, bucket, similarity and reliability),risk,input_schema,request_fit(how the request lines up with the input the tool needs),profile(the prediction), up to fouralternatives, thementionsits description names,routing_notes,first_steps,hintsandtimings_ms. - A stop:
toolis null andstopped_becausesays why, such as "the tools that fit best need authentication". Treat a stop as an answer, not an error.
The profile holds basis (where the numbers come from: the tool's own calls, its
bucket's calls, or nothing yet), success_probability, reliability,
expected_satisfaction_score from 0 to 3, and up to three likely_failure_modes with
probabilities.
select_tool(tool_id)
Looks up one tool, for example an alternative from a search. It returns the tool's details, risk and input
schema, its profile across all recorded calls, its mentions, the routing_edges its
description states (each with the exact excerpt), the other tools on its connector as
siblings, and hints. A tool that cannot be called says why in
excluded_because.
call_tool(tool_id, request, arguments, confirm_destructive=false)
Calls the tool and verifies the result. Pass the same request you searched with. It returns
result (the tool's own output), completed, tool_reported_error,
transport_error, prepared_by_search, next_steps, hints, and
an outcome:
| Field | Meaning |
|---|---|
satisfaction_level, satisfaction_score | How well the result answered the request, from 0 to 3 |
failure_mode | SATISFIED, or why not: EXECUTION_ERROR, EMPTY_RESULT, INCOMPLETE, INCORRECT, MISINTERPRETED, WRONG_TOOL or UNDERSPECIFIED_REQUEST |
confidence, needs_review | How sure the verdict is; needs_review is true when two judgments disagreed |
predicted_probability | The probability the profile gave the failure mode that happened |
prediction_matched, surprising | Whether the most likely predicted failure mode happened, and whether the outcome had a probability below 0.1 |
How to judge Hive fairly
Separate Hive's mistakes from everything else. Each step can fail on its own, and a connector's bug is not a routing bug.
| What happened | Count it as |
|---|---|
| Hive stopped on a request that needs a sign-in or that no callable connector serves | Correct, out of scope |
| Hive picked a callable tool that does the wrong job, when it should have stopped | Selection mistake |
| A better callable tool exists, in the alternatives or elsewhere | Selection mistake |
| Hive stopped although a callable tool clearly fits | Selection mistake |
| The tool errored, timed out, asked for payment, or returned bad data | Connector failure, unless Hive's verdict missed it |
The verdict disagrees with what you see in result | Verification mistake: the most useful kind to report |
The outcome was surprising while the profile's basis had many calls | Prediction mistake |
Hive refused a destructive call without confirm_destructive | Correct |
| A search took 1 to 4 seconds | Known, not a finding |
Hive sometimes returns the closest callable tool for a request it cannot serve. For example, "send an email
to [email protected]" picks AI-Mail's wait_for_email, which reads an inbox instead of sending.
Report that as a selection mistake. Do not then call the tool and also count its result against Hive.
Most profiles have few recorded calls, so basis is often bucket or empty and the
probabilities sit near their starting values. A wrong prediction on little evidence is expected; note the
basis when you report one.
Tips
- Read
hintsfirst. They name required inputs, confirmations and better alternatives. - Never invent arguments. Fill them only from the request or an earlier result. If a required value is missing, ask, or record that Hive flagged it.
- Search, then call within 10 minutes with the same request and tool. The call then reuses the search's prediction, so the outcome is judged against what you saw.
- Use
select_toolto compare an alternative or a mentioned tool before you call it. A tool id isnamespace/slug/name. - Do not retry failed calls blindly. Hive never retries a call, since a retry could repeat a change.
- Keep destructive calls out of tests unless the user agreed to them.
What to report
For each request you tested, report one line with these parts:
- The request, word for word.
- The tool you expected, or "none callable".
- Hive's pick or stop, with the
tool_id. - The outcome's
failure_modeandsatisfaction_score, if you called it. - Your judgment from the table above, and one sentence of evidence.
Finish with the number of requests in each judgment, and the mistakes you would fix first.