METHOD · 6 OCTOBER 2026
Bots and humans in the studies (exploratory): the method
Exploratory: post-hoc subgroups of the three finished studies, from a public-data heuristic with no ground truth. Every number counts as an extra trial of its parent study. The protocol was committed before any outcome was split by class.
Of the 323 accounts the studies measured, 22 look like bots, 151 like people using a web interface, 150 are unclear. For likely bots nothing robust can be said yet; for likely humans the protection results hold. All four studies.
The classifier
Its rules are in the protocol, fixed before any outcome was split. In short:
Bot evidence:
- client order ids on orders not routed through a builder (+2, or +1 above 10%);
- Ioc market orders off builder routes (+1);
- post-only-heavy order flow (+1);
- sub-second sequences of separate actions (+2, or +1);
- trading round the clock (+1);
- ≥ 100 actions per active day (+1);
- actions on the minute (+1);
- ≥ 90% of orders cancelled (+1).
Human evidence:
- the official web app’s market orders,
tif: "FrontendMarket"(−2, or −1); - a third-party front end, i.e. a builder fee on most fills (−1);
- a daily sleep gap (−1);
- under 180 filled orders in 180 days (−1).
Classes: score ≥ +2 is likely bot, ≤ −2 likely human, anything else unclear. Strict cut-off: ±3.
- The official web app sets no client order id. 1.5% of 123,638 FrontendMarket orders carry one.
- Third-party front ends do. Accounts trading mostly through a builder have a median cloid share of 59%, and they send Ioc market orders. Liquidated accounts trade mostly this way: a median 88% of their filled orders carry a builder fee.
Not usable:
- TWAP:
twapIdis null on every cached fill. - Approved API wallets:
extraAgentsis not in any cache, and fetching it needs the main session’s OK (below).
Coverage
| Population | n | Likely bot | Likely human | Unclear | Strict: bot / human / unclear |
|---|---|---|---|---|---|
| Compliance A, long-term winners (passed) | 20 | 1 (5%) | 8 (40%) | 11 (55%) | 0 / 3 / 17 |
| Compliance B, liquidated (passed) | 80 | 0 (0%, CI 0–5%) | 55 (69%) | 25 (31%) | 0 / 29 / 51 |
| Survival, cohort M | 212 | 19 (9%, CI 6–14%) | 84 (40%, CI 33–46%) | 109 (51%) | 10 / 32 / 170 |
| Guard defaults, test set | 63 | 8 (13%) | 24 (38%) | 31 (49%) | 4 / 12 / 47 |
| All study accounts (unique) | 323 | 22 (7%, CI 5–10%) | 151 (47%, CI 41–52%) | 150 (46%) | 11 / 66 / 246 |
- Insufficient data (under 10 orders): 7 accounts, all in B.
- Not in the caches: none.
- In M, the rules that fired:
- web-app market orders on ≥ 50% of entries: 42% of accounts;
- daily sleep gap: 31%;
- round the clock: 14%;
- API market orders: 7%;
- API client ids on ≥ 50%: 4%.
What the classes look like (all study accounts, medians):
| Likely bot | Likely human | Unclear | |
|---|---|---|---|
| Web-app market orders | 0% | 78% | 4% |
| API client ids | 15% | 0% | 0% |
| Ioc share of market orders | 93% | 0% | 0% |
| Actions per active day | 33 | 7 | 8 |
| Filled orders in the window | 2,635 | 366 | 536 |
Sanity check (descriptive)
The thresholds were set with these groups’ feature distributions in view, so agreement is weak support (protocol).
| Group | n | Bot | Human | Unclear |
|---|---|---|---|---|
| Market makers excluded from the studies (maker share ≥ 70%) | 119 | 23 | 5 | 91 (3 insufficient) |
| Public vaults (compliance cohort C, non-market-makers) | 11 | 3 | 3 | 5 |
| Low-activity accounts (M pre-screen, < 50 fills) | 39 | 1 | 0 | 38 (19 insufficient) |
- Mostly as expected, but the score is conservative.
- It rarely calls a market maker human (5 of 119), but it calls most of them unclear.
- Most of these market makers have only one cached page of fills and no order history, so the timing and order-type rules cannot fire.
- Disagreements:
- 3 of 11 public vaults score human (a sleep gap and web-app orders: vault leaders may trade by hand);
- 1 low-activity account scores bot.
- Low-activity accounts land in “unclear”, not “human”: a lack of data is not human evidence.
- Not checkable: HLP and other named market-making vaults are not in the caches.
Trial count
| Item | Measurements |
|---|---|
| Classifier: primary and strict cut-off | 2 |
| Compliance: 3 classes × (D2 per cohort, A − B gap) × 2 cut-offs | 12 |
| Survival: 3 classes × 4 measures × 2 runs × 2 cut-offs, plus 2 bot − human differences × 2 runs × 2 cut-offs | 56 |
| Guard defaults test: 3 classes × 4 measures × 2 cut-offs | 24 |
| Total | 94 |
- The parents’ counts rise by 12, 56 and 24.
- The Guard-defaults test set has now been looked at twice.
- Multiplicity: with 94 measurements, about five intervals would exclude zero by chance alone. That is one more reason to read every single-cell result here as exploratory.
- Not counted: the context numbers (return change, halts, feature medians) are reported as listed.
- Nothing else was tried: no rule, threshold or cut-off was changed after the protocol’s commit.
Caveats
- No ground truth. The score reads how orders were placed, not who decided.
- A person can run an API script, and an AI agent can drive a front end.
- One address can hold both, which is why much of the sample is unclear.
- Unclear is not a residual of noise. Accounts trading plain Gtc limit orders, with no client id and no web-app market orders, could be either. An SDK user who sets no client id and avoids Ioc looks like a manual limit-order trader.
- Agent wallets are invisible here.
extraAgentswas not fetched.- Even if it were, the official app’s one-click trading also signs with an agent, so an approved agent would not prove a bot.
- Whether the app’s agent can be told apart (for example by name) is untested.
- Builder routes are ambiguous. Wallet apps and Telegram front ends are mostly people; copy-trading and automation platforms also pay builders. H2 is weak for that reason.
- Recent history only: 2,000 orders and 10,000 fills. For active accounts, the order-type and timing features describe the last days to weeks.
- Small cells: 19 likely bots in M, 8 in the test set, 1 in the compliance cohorts.
- Post-hoc subgroups of studies whose headlines were already known; every parent study’s caveats apply (survivorship, what-if replay, main-dex perps, one 180-day period).