- 01Prompt engineeredEVIDENCE
- 02Premise challengedCRITIC
- 03Artifact executedBOUNDED
- 04Outcome verifiedPROVEN
AI evidence, with real people when it matters
Run the agents.
Prove the outcome.
TwoThumbs turns a raw ask into a defensible run: it engineers the brief, attacks weak premises, routes bounded execution across the right models, and verifies what actually happened. When agent evidence is not enough, add a small managed test with consented human participants.
3 real verdicts a day, no sign-up. Human studies are a separately scoped managed pilot.
https://twothumbs.co/mcpRemote MCP over HTTPS · durable async jobs · no local install
The loop
Fast enough to use.
Durable enough to trust.
TwoThumbs does not make you hold a request open while deep work runs. Queue one verdict or an atomic batch of up to 25, poll durable job IDs from MCP or CI, and get a terminal result with evidence. The same bounded runner and billing path handles every job, while explicit privacy controls keep developer-learning data separate from founder memory.
Adversarial planning
Challenge the premise before expensive execution begins.
Unattended runs
Submit, disconnect, and return to a durable job.
Evidence verdicts
Separate a plausible artifact from an outcome that actually worked.
Developer control
Opt out, export, or delete product telemetry from the same MCP.
Remote MCP / quick start
Put a verifier inside your workflow.
Connect the remote Streamable HTTP MCP once. Your agents and CI can submit single or batched verdicts, poll durable results, send structured feedback, and control their developer data without a local package or installer.
Run this in a terminal, then finish the client login flow:
codex mcp add twothumbs --url https://twothumbs.co/mcp && codex mcp login twothumbs
The command is the same in macOS Terminal, Linux shells, and Windows PowerShell 7 or Command Prompt.
Add the remote server to ~/.cursor/mcp.json on macOS/Linux, or %USERPROFILE%\.cursor\mcp.json on Windows:
{
"mcpServers": {
"twothumbs": {
"url": "https://twothumbs.co/mcp"
}
}
}Choose Streamable HTTP in an MCP-compatible client and use this HTTPS endpoint:
https://twothumbs.co/mcp
Add the server, then run /mcp inside Claude Code to authenticate:
claude mcp add --transport http --scope user twothumbs https://twothumbs.co/mcp
Commit this project-level .codex/config.toml, then inject TWOTHUMBS_API_KEY from your CI secret store:
[mcp_servers.twothumbs] url = "https://twothumbs.co/mcp" bearer_token_env_var = "TWOTHUMBS_API_KEY"
The endpoint is reachable worldwide over HTTPS where your client supports remote Streamable HTTP MCP and permits outbound access. Automated or paid use requires a valid key; network policy and client support still apply.
Have an access code? Add the MCP, then choose “Redeem a code” in its authorization screen or call redeem_code. Store the one-time returned key securely.
Managed human-testing pilot
Let real people find
what agents miss.
TwoThumbs can pair its automated evidence with a small, consented usability study. During this pilot, we scope the task, match participants, review submissions, and arrange payment manually. It is not yet a self-serve marketplace or a customer-facing human-testing MCP.
For product teams
Bring a public product surface and one decision you need help with. We will confirm the audience, task, participant reward, timing, evidence you will receive, and total price before anything begins.
Request a managed testFor participants
Join the pilot roster for possible paid usability tests. Matching is manual, selection is not guaranteed, and every invitation states the task, reward, review criteria, and payment timing before you choose whether to participate.
Apply to test productsClear boundary: roster consent and optional TwoThumbs dogfood invitations are separate. Each assigned study asks for its own consent, supports withdrawal, and does not turn participation into AI-training permission.
Multiple minds.
One accountable run.
The engineer + critic
Turns intent into completion criteria, retrieves relevant evidence, exposes contradictions, and sends consequential judgment through an adversarial model before execution.
Open the 60-second quickstart →The outcome verifier
Exercises the real surface and follows the work downstream. Mechanical success, visual comprehension, and the actual outcome remain separate gates.
Evidence, not self-scoring“Submitted” is not
the same as done.
Every visible affordance, across desktop and mobile.
Every form, with clearly marked synthetic data.
The row landed, the email arrived, the webhook fired.
TwoThumbs came from real bugs that sat behind polished interfaces for months—bugs AI code review never saw because the code looked plausible. The critic uses outcomes as evidence.
Keep verification
on speed dial.
Free
$03 surface scans every day. No signup.
Start with one line →100 credits
$9A small pack for active builds.
Request invoice →500 credits
$29For teams shipping on repeat.
Request invoice →