Researcher Dylan Castillo has tested whether major AI labs deliberately train models to generate images of pelicans riding bicycles. He ran 48 prompts across seven different models, including GPT-5.6 Terra, Claude Sonnet 5, and Gemini 3.5 Flash, repeating each test three times. Castillo used two other models to evaluate the outputs and found no evidence that these systems are specifically optimised for this task. The analysis showed that pelicans are not drawn better than other animals and bicycles are not drawn better than other vehicles. No lab demonstrated a significant advantage on this specific combination compared to its general performance.
This investigation matters because it challenges the assumption that AI companies secretly optimise for obscure tasks to boost benchmark scores. The lack of improvement suggests current evaluation metrics may not detect hidden training priorities. If labs were indeed ‘pelicanmaxxing’, users would see a distinct edge in these specific scenarios. The data indicates such an edge does not exist in practice.
- No model showed better pelican or bicycle rendering than expected.
- GLM-5.2 showed a small, non-significant boost on the specific cell.
- Scenes did not appear memorised or artificially inflated.




