AI evidence, with real people when it matters

Run the agents.
Prove the outcome.

TwoThumbs turns a raw ask into a defensible run: it engineers the brief, attacks weak premises, routes bounded execution across the right models, and verifies what actually happened. When agent evidence is not enough, add a small managed test with consented human participants.

Works withClaude Code · Codex · Cursor · CI / headless

3 real verdicts a day, no sign-up. Human studies are a separately scoped managed pilot.

https://twothumbs.co/mcp

Remote MCP over HTTPS · durable async jobs · no local install

Contact form / 0042● ●
  • 01Prompt engineeredEVIDENCE
  • 02Premise challengedCRITIC
  • 03Artifact executedBOUNDED
  • 04Outcome verifiedPROVEN
Ship it

The loop

Fast enough to use.
Durable enough to trust.

TwoThumbs does not make you hold a request open while deep work runs. Queue one verdict or an atomic batch of up to 25, poll durable job IDs from MCP or CI, and get a terminal result with evidence. The same bounded runner and billing path handles every job, while explicit privacy controls keep developer-learning data separate from founder memory.

01

Adversarial planning

Challenge the premise before expensive execution begins.

02

Unattended runs

Submit, disconnect, and return to a durable job.

03

Evidence verdicts

Separate a plausible artifact from an outcome that actually worked.

04

Developer control

Opt out, export, or delete product telemetry from the same MCP.

Remote MCP / quick start

Put a verifier inside your workflow.

Connect the remote Streamable HTTP MCP once. Your agents and CI can submit single or batched verdicts, poll durable results, send structured feedback, and control their developer data without a local package or installer.

Choose Streamable HTTP in an MCP-compatible client and use this HTTPS endpoint:

https://twothumbs.co/mcp

The endpoint is reachable worldwide over HTTPS where your client supports remote Streamable HTTP MCP and permits outbound access. Automated or paid use requires a valid key; network policy and client support still apply.

Have an access code? Add the MCP, then choose “Redeem a code” in its authorization screen or call redeem_code. Store the one-time returned key securely.

All setup and CI options →

Managed human-testing pilot

Let real people find
what agents miss.

TwoThumbs can pair its automated evidence with a small, consented usability study. During this pilot, we scope the task, match participants, review submissions, and arrange payment manually. It is not yet a self-serve marketplace or a customer-facing human-testing MCP.

01

For product teams

Bring a public product surface and one decision you need help with. We will confirm the audience, task, participant reward, timing, evidence you will receive, and total price before anything begins.

Request a managed test
02

For participants

Join the pilot roster for possible paid usability tests. Matching is manual, selection is not guaranteed, and every invitation states the task, reward, review criteria, and payment timing before you choose whether to participate.

Apply to test products

Clear boundary: roster consent and optional TwoThumbs dogfood invitations are separate. Each assigned study asks for its own consent, supports withdrawal, and does not turn participation into AI-training permission.

Multiple minds.
One accountable run.

01

The engineer + critic

Turns intent into completion criteria, retrieves relevant evidence, exposes contradictions, and sends consequential judgment through an adversarial model before execution.

Open the 60-second quickstart →
02

The outcome verifier

Exercises the real surface and follows the work downstream. Mechanical success, visual comprehension, and the actual outcome remain separate gates.

Evidence, not self-scoring

“Submitted” is not
the same as done.

Click

Every visible affordance, across desktop and mobile.

Submit

Every form, with clearly marked synthetic data.

Prove

The row landed, the email arrived, the webhook fired.

TwoThumbs came from real bugs that sat behind polished interfaces for months—bugs AI code review never saw because the code looked plausible. The critic uses outcomes as evidence.

Keep verification
on speed dial.