Show HN: Pelican-bicycle alternatives

87 points
1/21/1970
7 hours ago
by tkgally

Comments


svcrunch

I'd like to mention the Little Dorrit Benchmark [1] which I have been running for a couple of years now. It has a few nice features:

1. It tests visual reasoning and structured output in a single task.

2. It seems to sort correctly on advancing general intelligence. As a counterexample, if I'm not misremembering, artificialanalysis.ai made some changes to their benchmark recently after Astra ranked below several older models.

3. While models have gotten significantly better in the past 2 years, the top model is still at 0.78 F1, so the test is not yet saturated. As a reference point, when I started, the top models were in the [0.1, 0.2] range.

[1] https://dorrit.pairsys.ai/

5 hours ago

pohl

Gemini 2.5 Pro is the only model with a sense of where a ferris wheel operator would be.

7 minutes ago

water-drummer

Ok the Grok ones are cute

3 minutes ago

vova_hn2

Website looks very cool, Fable's octopus-organist looks very cute, but I feel like this benchmark (generate an SVG by a short and slightly ridiculous description) in general has been completely Goodharted [0].

I think they all just added a bunch of similar tasks to their training sets, so we cannot judge true emergent capabilities of the models anymore.

[0] https://en.wikipedia.org/wiki/Goodhart%27s_law

5 hours ago

andy_ppp

“I think they all just added a bunch of similar tasks to their training sets, so we cannot judge true emergent capabilities of the models anymore.”

This has been the plan since the start of all this, they regurgitate code in ever better forms but they still aren’t inventing new things yet.

an hour ago

GaggiX

I don't think a bunch of similar tasks can really saturate the "create a SVG of X", because the model should have a quite good spatial understanding of the world and how everything interacts.

For example Gemini 3.8 Flash seems very impressive at first glance but the results are not actually very coherent, this shows that its "world model" is not particularly great (compare to SOTA models).

4 hours ago

samayashar

All models are pretty good now at generating these images. Back in the day, I remember experimenting with the pelican images and most of the models couldn't align the legs with the wheels. Right now as well, GPT messed up an octopus leg by originating it through the instrument rather than the octopus itself.

I think that intertwining two entities (living/non-living) is still challenging but overall they're pretty sound.

6 hours ago

honeycrispy

> All models are pretty good now at generating these images.

That's pretty generous.

4 hours ago

CamperBob2

All models are pretty good now at generating these images.

Not zebras. If you want to see how bad SVG output still is, ask for a zebra riding a scooter.

5 hours ago

outlore

Anyone else surprised the generations look so remarkably similar? All of these models have “independently” generalized that the moose should roughly be standing at the same position (left) or that the giraffe should have a certain color palette.

2 hours ago

Terr_

With respect to bikes (with or without pelicans) there's a strong natural bias because people displaying bikes tend to want to show off the side with the gears.

More-generally, I suspect an influence from how left-to-right languages (i.e. English) affect comic layouts. Overcoming that bias often means using vertical space to exploit the top-to-bottom habit instead. (Consider the rarity of an English-language comic panel where action is from bottom-right to top-left.)

2 hours ago

simonw

I love these.

13 minutes ago

andy_ppp

Love the output from Qwen 3.8 it seems very impressive for the cost! Why does Gemini 3.8 flash blur everything? What are Google playing at!

an hour ago

steinvakt2

Feels like google has a different training set than the others?

6 hours ago

kennywinker

Would love to see Qwen3.8-27b here, since that is the model most people are running locally.

3 hours ago

zaphar

I notice none of the octopi seem to be actually facing the organ.

3 hours ago

sajithdilshan

Interesting, out of all examples Gemini 3.8 is the best for me. Also the image style is different and more vibrant than others

5 hours ago

Jordan-117

Better than Astra and Fable? It looks quite pretty and even impressive at times if you squint, but look closer and it falls apart in terms of coherency. And I say that as somebody who mains Gemini 3.8.

4 hours ago

eddytrex_

Does a test of instructions how to fold origami figures in a SVG/jpeg exist? Or could be useful?

5 hours ago

BrokenCogs

Gemini 3.8 flash seems to (subjectively) be the outlier in terms of performance to cost ratio?

4 hours ago

simonw

Google Gemini are the only team who have openly had staff deliberately spend time on SVG performance: https://twitter.com/sunjiao123sun_/status/202455551655137292...

See also this Jeff Dean tweet showing off their animals-in-vehicles abilities: https://twitter.com/JeffDean/status/2024525132266688757

11 minutes ago

gpt5

It’s very likely they all add svg generation into the training data. It’s part of the reason it’s no longer a good benchmark (unless you need to generate SVGs).

4 minutes ago

pixelesque

Its giraffe / grandfather clock one is pretty bad... (two necks? wearing a suit?)

Weird as well, it's clearly pulled out some 1884 patent on clock designs, and a quick ddg/google doesn't show it as anything to do with grandfather clocks.

3 hours ago

GaggiX

Gemini 3.8 Flash results are often not very coherent but it does put a lot of shading and details to hide the fact.

4 hours ago

sceptic123

> An elephant typing on a typewriter

A monkey, surely?

5 hours ago

neilellis

Well that benchmark is now saturated, what next. How fast you can hack the pentagon?

4 hours ago

dustfinger

It is interesting how similar the designs are across the models.

4 hours ago

qiine

Asking to animate it add an interesting layer of difficulty

5 hours ago

ormax3

I noticed in the "A penguin juggling chainsaws" prompt that Qwen created an animated svg

4 hours ago

input_sh

I'd say at least half of Qwen's 2026 runs are animated.

The only other one I've spotted is animated is Gemini 3.0's 2025 run of an elephant.

3 hours ago

mock-possum

Try asking an LLM to draw you the cool S.

3 hours ago

amysox

Or something like, "Draw an S, then a more different S, close it up real good here, then using consummate V's, add teeth, and scales, and eyebrows, and legs. And then add smoke, and fire, and some wings, and one of those big beefy arms for good measure." :D :D :D

an hour ago

dcreater

Why is this a good test?

4 hours ago

simonw

Because it's one of the few ways of comparing models that lets you instantly evaluate them visually. That makes it more comprehensible than a numeric score on a benchmark.

10 minutes ago

villish

The 3 US models have their own style.

Qwen3.8 is very clearly distilled from Claude models.

4 hours ago