TESTS / 02 ENTRIES

Model tests without the mystery.

Same inputs when possible. Attempts and fixes disclosed. The original post linked directly.

Grok 4.6 vs Opus 5

The same one-shot prompt, one attempt each, and no fixes. I chose the winner after playing both games.

VIEW TEST ON X
Prompt
Identical
Attempts
One each
Fixes
Zero
Winner
Grok 4.6
TAKEAWAY

Opus looked better first. Grok built the game I preferred playing.

mona-lisa-1 vs GPT Image

A mystery Arena model and GPT Image received the same prompt. The outputs were shown side by side.

VIEW TEST ON X
Prompt
Identical
Format
Side by side
Model
mona-lisa-1
Compared with
GPT Image
TAKEAWAY

The mystery model held its own against a known image model.

THE METHOD

Four rules for every comparison.

  1. 01Use the same inputs when possible.
  2. 02Disclose attempts, edits, and settings.
  3. 03Show both outputs.
  4. 04Separate facts from personal taste.