InternalAI pipeline eval

How right is the AI?

8 recordings, 623 clips with a human verdict. Last experiment logged Sep 11, 2026.

By recording

  1. Pickup 2v2, 29 Mar

    Pickup

    Mar 29, 2026 · 55:03 of footage · 367 candidates

    Open review
    238 / 367 reviewed65%
    Real shots among proposals
    62%n=235
    Made/miss accuracy
    91%n=145

    made as positive:Precision 86%Recall 93%F1 90%

    Shooter agreement
    not yet
    Review speed
    not yet
  2. Solo shooting, 15 June

    Solo

    Jun 15, 2026 · 12:21 of footage · 75 candidates

    Open review
    75 / 75 reviewed100%
    Real shots among proposals
    100%n=75
    Made/miss accuracy
    93%n=75

    made as positive:Precision 86%Recall 100%F1 93%

    Shooter agreement
    not yet
    Review speed
    not yet
  3. Solo shooting, 3 June

    Solo

    Jun 3, 2026 · 31:17 of footage · 238 candidates

    Open review
    178 / 238 reviewed75%
    Real shots among proposals
    84%n=178
    Made/miss accuracy
    90%n=149

    made as positive:Precision 80%Recall 96%F1 87%

    Shooter agreement
    not yet
    Review speed
    6.7sper shot
  4. 2v2, Hunters Hill, 18 Jan

    Small-sided

    Jan 18, 2026 · 18:44 of footage · 83 candidates · 4 filtered

    Open review
    83 / 83 reviewed100%
    Real shots among proposals
    77%n=83
    Made/miss accuracy
    86%n=64

    made as positive:Precision 76%Recall 92%F1 83%

    Shooter agreement
    96%n=50
    Review speed
    8.2sper shot

Over time

226 logged runs from the experiment ledger. Tap a dot for the method and cost.

  • Eltham Rnd 1
  • game_hd 2v2
  • Solo shooting, 15 June
  • other

Shot recall

30 runs
0%50%100%AprJunAug96.6%

Shot precision

30 runs
0%50%100%AprJunAug53.8%

Shot F1

29 runs
0%50%100%AprJunAug69.1%

Accuracy

86 runs
0%50%100%AprJunAug100.0%

Made/miss accuracy

1 run
0%50%100%AprJunAug76.9%

Candidate recall (5s)

18 runs
0%50%100%AprJunAug0.0%

CV-A accuracy

30 runs
0%50%100%AprJunAug82.8%

Made/miss AUC

2 runs
0%50%100%AprJunAug66.4%

How to read this

A clip's human verdict is the strict majority of its reviewers; ties count as split and drop out of every rate. Shot precision is the share of reviewed candidates that were real shots. Made/miss numbers compare the AI's call with that verdict, with "made" as the positive class, so precision is "when the AI says made, was it" and recall is "of the real makes, how many did it call". Recall of shots themselves stays unmeasured until the audit recording is labelled.

Get early access to GameTape
Join Waitlist