Quasa
Use QUASA App
Join the pioneer of Web3 crypto freelancing today!
Open
Technology

VLMs Can Already Hunt for the “Interesting.” But They’re Still Bad at Walking Away from What They’ve Found

|Author: Viacheslav Vasipenok|5 min read| 38
VLMs Can Already Hunt for the “Interesting.” But They’re Still Bad at Walking Away from What They’ve Found

In 2008, a website called PicBreeder quietly demonstrated one of the most powerful ideas in artificial creativity. There was no target image. No fitness function. No instructions about what to create.

VLMs Can Already Hunt for the “Interesting.” But They’re Still Bad at Walking Away from What They’ve FoundThousands of people simply browsed grids of evolving abstract images, picked the ones that felt promising or intriguing, and “bred” them — combining and mutating the underlying neural networks (CPPNs) to produce offspring. 

Over generations, starting from pure noise, recognizable forms emerged: faces, animals, vehicles, skulls, insects, and strange hybrid creatures no one had explicitly asked for. The process was open-ended discovery in its purest form.

Kenneth Stanley captured the deeper lesson in his book Why Greatness Cannot Be Planned: the most interesting discoveries often come not from chasing a predefined goal, but from following curiosity and preserving unexpected novelty.

Now, researchers from Sakana AI, MIT, and NYU have asked a natural next question: Can modern vision-language models (VLMs) play the same game?


The Experiment: PicBreeder Without Humans

VLMs Can Already Hunt for the “Interesting.” But They’re Still Bad at Walking Away from What They’ve FoundIn their new work (“In Search of the Ingredients of Open-Endedness: Replicating PicBreeder with Large Vision-Language Models”), accepted at GECCO 2026, the team built a fully automated version of PicBreeder driven entirely by frontier VLMs (such as Gemini 1.5 Pro).

The setup deliberately mirrored the original:

  • A shared, growing archive of images.
  • Multiple VLM “agents” that could view the archive.
  • Agents chose images they found interesting.
  • They bred new variations (through mutation and crossover of the underlying networks).
  • They published their favorites back into the archive.
  • Other agents could evaluate and rate the work.

Crucially, no target image was ever provided. There was no progress metric, no reward for reaching a specific concept, and no external goal. The only pressure was the agents’ own internal sense of what counted as “interesting.”


What Worked

VLMs Can Already Hunt for the “Interesting.” But They’re Still Bad at Walking Away from What They’ve FoundThe VLM agents were surprisingly capable at the first part of the task. They could identify visual and semantic hooks in the images — eyes, symmetry, organic shapes, mechanical forms — and use those as starting points for further evolution. When the researchers introduced multiple agents with different “personalities” (subtle variations in their system prompts, e.g., “curious,” “minimalist,” “surreal”), the archive became noticeably richer and more diverse. In terms of semantic coverage and recall of recognizable concepts, diversity among agents helped close some of the gap with the historical human-generated PicBreeder archive.

This suggests that population-level diversity — having many agents with slightly different tastes — is a powerful lever for open-ended exploration, much like having thousands of human users with varied interests.


The Core Limitation: Premature Fixation

VLMs Can Already Hunt for the “Interesting.” But They’re Still Bad at Walking Away from What They’ve FoundHowever, a clear qualitative gap remained. While humans in the original PicBreeder could spot a random, seemingly useless strangeness and pivot an entire lineage toward it, the VLM agents tended to do the opposite.

Once they discovered something even moderately coherent — a duck-like shape, a skull, a pair of legs, a symmetrical motif — they would often double down on it. Instead of using the novelty as a springboard to explore new territory, they would refine, polish, and exploit the same motif across many generations. The evolutionary trees became unbalanced: a few “successful” parents dominated, while promising but less immediately attractive branches were neglected or abandoned.

Humans appear better at two crucial skills the current VLMs lack:

  • Changing their own criterion of interestingness mid-process.
  • Preserving and amplifying weak signals — strange mutations that don’t yet look promising but might lead somewhere new.

The agents were good at recognizing novelty when it appeared in front of them, but poor at actively seeking it by abandoning locally attractive attractors.


Why This Matters

This isn’t just an interesting experiment in evolutionary art. It speaks directly to a fundamental challenge in building more capable AI systems: open-endedness.

VLMs Can Already Hunt for the “Interesting.” But They’re Still Bad at Walking Away from What They’ve FoundMost current AI training is heavily goal-directed — optimize a loss, follow instructions, maximize a reward. PicBreeder-style exploration offers a different paradigm: discovery through curiosity-driven, goal-less iteration. If we want AI that can generate genuinely new ideas, scientific hypotheses, or cultural artifacts rather than just remix existing patterns, we need systems that can do more than chase the next obvious improvement.

The Sakana/MIT/NYU results show that today’s VLMs already possess some of the ingredients:

  • Visual-semantic understanding;
  • Ability to evaluate and select;
  • Capacity for iterative generation.

But they are still missing key pieces for true open-ended discovery:

  • Mechanisms to dynamically shift their own notion of “interesting”;
  • Better handling of long-term exploration vs. short-term exploitation;
  • Ways to maintain diversity across an entire population of agents without heavy orchestration.

The Path Forward

VLMs Can Already Hunt for the “Interesting.” But They’re Still Bad at Walking Away from What They’ve FoundThe researchers found that simple interventions helped: giving agents distinct personalities, limiting how much past context they see (to avoid getting stuck in loops), and occasionally injecting noise. But these are patches. The deeper lesson is that recognizing novelty is not enough. An agent must also be willing to leave a promising direction when better ones might lie elsewhere — and to keep weak, half-formed ideas alive long enough for them to mature.

PicBreeder worked because humans are remarkably good at exactly that kind of flexible, serendipity-driven search. Replicating that capability in artificial systems remains one of the most exciting open problems in AI.

The full technical blog post and interactive demo are available at pub.sakana.ai/picbreeder-vlm. The accompanying paper and dataset have also been released.

The original PicBreeder showed us what collective human curiosity can achieve when freed from objectives. The new VLM version shows us how far we still have to go — and gives us a concrete benchmark for measuring progress toward genuinely open-ended artificial discovery.

---

VLMs Can Already Hunt for the “Interesting.” But They’re Still Bad at Walking Away from What They’ve FoundAlso read:

---

Thank you!

Share:

Subscribe to our newsletter

Get the latest Web3, AI, and crypto news delivered straight to your inbox.

0