When you aren't sure if an LLM can write an svg well, or that it will be able to form a pelican shape, or animate a bicycle, it's a good test. After that, it's all judgement: how detailed should the pelican be? pelicans are the wrong shape for a bicycle by default, so how much can I change its physiology to match using a bicycle before it isn't a pelican? Do I care about how well the client is able to render complex geometry?
It's not that there isn't room to do better, or that it doesn't tell you anything at all, but rather we've reached a point where what it tells us isn't very clear anymore.
https://www.booooooom.com/2016/05/09/bicycles-built-based-on...
Totally agree though, anyone with a vague understanding of how bikes works ignores the pelican because they know the bike is unrideable in the first place.
A good new benchmark should have awful performance to start and there should be a lot of headroom for improvement. This benchmark is also intentionally difficult and requires the LLM to develop the animation through spatial reasoning and first principals rather than existing video generation pipelines. Similar to how SVG generation was out of distribution for most models a year ago.
Imho we don’t need to make benchmarks that draw the whole 3D world. Pelican’s drawing is really nice in its simplicity and complexity at the same time.
It seems to be an obscene waste of compute time to generate useless 3D worlds that are just a bragging - 3D is really heavy discipline to make it right, see Mark Zuckerberg’s ceased attempt with 3D VR…
Multiply it by thousands times as a lot of people have found out threejs lib and prompt “generate 3D world and make no mistake” are new orange/black.
That took a fair amount of custom tuning and I had to create a tuning view to get some of the behaviors right.
But it was enough fun that I generalized it to take in ~any scene description from a film. It goes out and gets more detailed descriptions and film stills if available but also takes custom stills if you provide them.
My test scene was the Gauntlet scene from Apocalypto. It is low fidelity but does a pretty amazing sequence with somewhat believable physics of the javelins etc.
Here is the docs page with the vertical takeoff / 88 miles an hour time travel: https://contextify.sh/docs
I can share some of the Apocalypto bit if anyone is interested.
That page with the time machine may not make it obvious, but the product does have a linux client!
It doesn't have the same app window and summarization of the macos version, but the transcript ingestion engine is efficient and the real value is in leveraging the database it builds using the packaged `/total-recall` or your own use of the api. (which is not yet documented but discoverable)
I am building the windows version now. I have been for the past four days. It uses a shared swift-core with the macOS and Linux versions which has been part of the reason it has taken "so long."
If anyone is on windows (or linux!) and would be willing to try it either of the clients my email is in my profile, I would be grateful.
Also I'm working up a short post with that apocolypto anim now.
When Fable was first released the day-1 demos of it on Twitter (presumably from people who were given early access, and/or Anthropic employees) were pretty much 100% three.js stuff. Yes, it looks nice, but it doesn't tell me any better than an Erdos proof whether the LLM will be able to run my vending machine.
Are you sure about that? Presumably the spatial concept of "inside" and the computer memory concept of "inside" have different contextual embeddings in an LLM, much in the same way that neural activation patterns vary in the human brain when we reason about different domains.
This weekend I've been converting a game from three.js to ogl.js in order to see if I can optimise the time-to-interactive loading time. I took the three.js driven page weight from about 600KB (500KB being three.js) to about 50KB, and reduced the loading time from multiple seconds on a 4G mobile connection to around 0.5s.
This has mostly been a combination of Opus 5 and Sonnet 5 in Claude Code. It very clearly has a good grasp of WebGl, and of what impacts page loading times and rendering speed. It was able to drive Claude Code's integrated browser to measure the impact of changes, and as I spiked out a test of ogl.js it could test the differences changes made.
It's not the best game (https://tinyslots.ooer.com) but that's on me. As an exercise in building 3D in a browser, and in page speed optimization, with Claude models I am really impressed.
This is what you get. It's like the more general concept of starting with the written word (or 1000 words) and then replacing it with a picture. You've done something strikingly different, but is it serving the same function?
It's fascinating to see this stuff combine such disparate sources in unexpected ways. But it is parrot, just not in the way you're expecting.
a real benchmark is instead running evals on your own traces, and building a cost/quality/speed profile for models based on real workloads. but it doesn't get you a shiny video you can post on twitter.
Definitely impressive demo, I do wonder though if the countless artwork, films, images etc produced over many decades around Lord of the Rings somewhat influenced the outcome of this though.
"Getting Started with Google Wave": https://www.youtube.com/watch?v=eKUAqNGVwX0
Years later Apache moved it to read only because of low community activity.
The archived git repo on GitHub remains available to clone and revive as a fork.
wave failed for weird google organizational reasons far more than anything inherent to the product or tech
"Draw an animation of this long ass scene from a movie, and only call me when everything works e2e" can be.
Consuming radium and using uranium glass, that’s what we’re doing.
I think LLMs will be excellent glue of "find the right function/button and run/push it" but design without constraints and they just explode immediately
Why do you say this? I have some experience of manufacturing processes and am not seeing where AI would be useful other than to drive the robots which we can already do quite well without AI (see all the dark/lights-out factories that already exist).
Where do you see it being useful? An example would be nice.
Asking AI to design real world objects doesn't work very well because all of its tests involve proxies and thus miss things that are glaringly obvious when the object is actually built.
We already have sneaker designs and the equipment to manufacture them. Whatever it spits out is going to be, at best, a mediocre clone of something that already exists. What exactly is the point?
I’ve been trying it on them all and can’t find one that does it consistently. The best will tell me they can’t. The worst confidently point out one of countless Waldo-likes.
I gave Fable a jpeg and asked to draw an SVG, using a loop that renders the SVG into an image so Fable can inspect it.
Results looked like drawing of a 5 year old.
I wonder whether we are entering the era of throwaway software. Just like cheap plastics and improved processes has enabled us to rapidly manufacture anything we want for a very low price, maybe LLMs give us the same for software. Produce it cheaply and if it breaks throws it away and reproduce it.
On the one hand yes, almost every task I work at now is one off one of scripts I throw away.
The question is - where does the software the spec or the “code”.
A really complex game will probably always be token heavy. At least for the next few years code is still not free.
But certain software is just iterative by design. If we mean we regenerate all the for loops of a game from scratch, sure but I think “code” Is really more spec then implementation, and we’ll want to continue building things through iteration.
And even on the for loop point - Do you really want to spend millions of tokens rewriting a game every time you need to make balance changes?
Why would anyone still use off-the-shelf software when they can have a system that has access to all data, can transform it into any form, and can export it in any format?
After years of thinking that I needed to develop a decent movie management system for my own films or a columnar browser for large CSV files, Claude and Qwen each delivered exactly what I needed in just a day.
Have they? Most of the world production is tied down to expensive factories and machines. Yes, we have more products, but that the result of the global trade, which is a very complex system.
> Produce it cheaply and if it breaks throws it away and reproduce it.
I don't know why everyone would ever wants this. It's been parroted since forever, but the true usefulness of software is to be able to build it once and runs it indefinitely. If some edge case occurs, I fix it. Which is way cheaper than rebuilding the whole thing. The goal is to have something like OpenBSD's ed[0] or dmesg[1], which you only touch every few years or so
[0] https://github.com/openbsd/src/commits/master/bin/ed
[1] https://github.com/openbsd/src/commits/master/sbin/dmesg/dme...
Bilbo's house is actually described in The Hobbit and the exterior isn't really described at all in Fellowship. The prologue of Fellowship (Concerning Hobbits) mentions hobbits like round doors and windows and the fact some hobbit homes are underground, but the turf-dome design here is not mentioned. It actually mentions hobbit homes typically have bulging walls, so unless you've read the The Hobbit, you might not picture this entirely-underground style.
In The Hobbit, his home is described as a (nice) hole in "The Hill" with a perfectly round front door and round windows, which could imply the design here.
However I do agree that the results are much worse when you try to use them. They look great in screenshots and video clips which makes them perfect for content farmers.
All of the LLM generated games I’ve played have been really bad to play, though. I even tried my hand at a simple game, thinking I could iterate on it with prompts to fix some rough edges. After the initial productivity burst every change turned into a slog of tokens with one thing changing and something else breaking it. I would try to use my remaining weekly token budget across Anthropic and OpenAI to refine it at the end of every week but after a couple weeks it felt like I wouldn’t be getting anywhere without scrapping it and going back to having the LLM build it one step at a time with my careful instruction.
Which, in retrospect, is the only way I can get usable output of an LLM for anything complicated, so it’s not surprising. It’s a fun reality check project though.
On the one hand I think most of us are incredibly impressed because we know, that quick demo would have taken us months of work to build in the before times.
On the other hand the promise is a cure for cancer and the end of all work.
So when everyone is telling you “skill issue is why you can’t one shot WoW”. It’s hard to know how you’re supposed to feel about Karpathy advertising one shot custom virtual worlds but giving you slop. Incredibly impressive slop when compared to how long it would take to create it just 3 years ago, not so much compared to Elon saying - “by the end of this year, grok will create a version of the odyssey that competes with Nolan’s”
These conversations get frustrating since one group is saying the dancing is bad, a second group is saying it's good (for a bear), and a third group is saying we are a few years from a bear-only dancing industry.
I guess this is the average story and, similarly, the average game is boring and predictable
Early generative AI at least had the virtue of relentless, unsettling weirdness, in the same way that generative art from the late 90s and early 2000s did. A handful of people made creative use of that spooky weirdness.
Now it turns out "Airspace" art.
If you took the best, most creative, human writer in the world, and for thought experiment reasons they had amnesia (to mimic AI blank context windows) specifically while you asked them for a story idea 100 times in a row, my expectation is that this human would also give you the same idea at least 80 times out of that 100.
* still better than the mean human, but even the top 0.1% of humans aren't all professional authors.
What you think is better is not what I think is better. Imagination is not storytelling. You ask the "mean human" to write, it's going to be worse than an LLM in spelling and grammar, if you get anything at all.
That's oddly specific and it'd really hurt my cousins feelings :)
I guess the next question is if they can be made fun without too much additional work with a human guiding the AI.
They can't. If you think about how these things are trained it's blatantly obvious fun is an impossible metric to optimize them for
..feels like there was a black mirror episode about that though
Used to be if a game looked that good it probably had time spent on the game part too.
Whenever I think about AI games I find myself thinking about Tiny Wings, Flappy Bird and Angry Birds. Three simple, elegant games.
It is easy to see what makes Tiny Wings so completely loveable — it has a sculpted, adorable, perfect charm with a cleverly inverted game mechanic that has a calibrated level of exasperation and reward.
But why were Flappy Bird and Angry Birds, very basic games with very old game mechanics, so charming?
It seems equally impossible to imagine an AI coming up with a game with the quality of any of them, even with maximised creativity. But explaining why for Flappy Bird seems quite difficult, especially when you consider it uses some stolen visuals!
Maybe it's the smoothness of motion that makes these games understandable and LLMs seem to consistently fail at that. Ask them to do something snowboarding and they go really hard on the physics since it seems like they don't know what kind of approxmations feel good.
(It's been a while since I was in the game industry, so IDK quite how accurate this is).
If I see one more “one shot MMO” where you just walk around and do absolutey nothing or another menu slop idle battler or rogulike deck builder I’m going to go Postal in Minecraft.
These one-shot products aren't games. They're barely even demos. I don't even know what to call them. For a mature framework like Phaser to sell-out like this and create a vibecoded platform for vibecoded games is shocking.
Computer graphics will have enormous applications because they are directly controllable by LLM-generated code. Video models are probabilistic and less suitable when precision matters. In education, for example, we need exact visuals. If an AI wants to plot y = sin(x), it should generate the precise graph through computer graphics rather than approximate it with a video model.
What I find funny is that approximate to computer graphics are video games. When I ask AI about a decision available to me in a video game AI completely fails, OFTEN. I assume all the forums and changes made to a game over time might be quite confusing for AI. But I've also seen it completely make up characters and decisions and weapons and so on about some very clearly defined games and paths. It's an interesting dynamic.
It always brings to my mind some words from Rich Hickey:
I think we’re in this world I’d like to call “guardrail programming”. It’s really sad: we’re like, “I can make change because I have tests!”. Who does that? Who drives their car around, banging against the guardrails, saying “whoah, I’m so glad I have these guardrails so I can make it to the show on time!”
I don’t think I really have a point to make here, other than it just feels like someone’s released a bunch of carnival bumper cars onto the highways.Not difficult to see why the employees of AI firms are thrilled with it though, eh?
I like the analogy.
I guess this is why we got this before cars sold without steering wheels: literal guardrails on the literal roads are somewhat more expensive, especially for the people who keep bouncing off them on the way to their destination.
Also, where the guardrails are absent: oh look, felonies. https://www.google.com/search?q=ai+hacks+company&tbm=nws
As with painting, after a while there's nothing really new to paint, we genuinely need 0 new software. We need to fix our broken physical world, our social lives, our kids and what's left of our democracies.
This software crap is done, leave it to the nerds.
Textbook and blackboard > ipad.
This is an odd take, given that Karpathy is certainly aware that the LotR films absolutely did create Bag End in digital format; that their creation was outstandingly high quality; and that Claude’s output here very obviously “leans heavily” on their prior art.
I wonder if Flash is still popular... LLM can use that instead...?
[0] Seriously. Get used to mentally prefixing his and Boris Cherny's name like this, every time you see them quoted. These people are speaking while employed; there is no chance they are not aligned with the employers who will make them wealthy. The tech industry does like to pretend that for some reason AI people, uniquely, speak thoughts unbiased and for themselves or even for science or humanity.
that's why these things are actually pretty good at openscad/freecad/F360 mcps , the visual reality is enforced and guaranteed by rigor in the interpretation engine that is anchored to human physical reality.
There are people in their right mind who would do that and their are already examples of people who did similar things.
But maybe not in the future if people would confuse all the effort with AI
I think it's interesting that the "Bag's End" interpretation in the video clearly looks like the one from the movies, but generated here as a three.js 3D asset.
It makes sense that the movies (or shots/frames from them) were in the training data, and I can also easily imagine an association in concept space between the textual description of Bag's End and the frames from the movie.
But how on earth does the model then go on and convert the latent representation of those images into coordinates for a 3D mesh, without ever even restoring the image? In what kind of representation are the images from the movies stored that it can do that?
I like where Karpathy is going; I had the same thoughts about LLM generated slop scenery. I just want some variety of scenery for the goblins to get massacred in in whatever fantasy slop game I play.
8 months ago, he was (very reasonably) claiming that reliable agents are at least a decade away, but this now goes against the interest of his employer, so the narrative has been changed.
Or maybe over these 8 months agents improved a lot? You know, few years ago many AI experts predicted that things we are routinely doing now with AI are decades away. I mean how can you look at this post and not be impressed? It's insane what AI is currently capable of.
Except it's not ~free, it cost ~$10. And no one in their right mind would ever exchange $10 for that crap output except in this brief moment that we're in because it's fun and surprising to see what will happen. The actual result is as close to useless as it's possible to be - it's not interesting in its own right, it's not aesthetically pleasing, nor funny, nor informative. It's just slop.
It’s a useful benchmark (aside from being “cute”) because of its simplicity, both in how many output tokens it takes (though I understand some models think a lot now to do it) and how easily one can subjectively judge. It’s this efficient as a benchmark of performance.
Making a long video takes way more tokens, and presumably is a lot tougher to easily compare. swillison has a presentation that’s pelicans from 2023-present (roughly) showing the progression. Imagine “lord of the rings videos from 2026-2029” or whatever, it would take a long time to watch and be harder to judge, and probably just end up being a comparison of screenshots anyway.
TLDR I feel like the post misunderstands the role of the pelican thing though if find it very hard to believe he really doesn’t understand, so maybe I’m missing something.
There’s a reason AI slop games have literally zero engagement. Last summer that stupid flying game blew up. Maybe a million people “played” the game. Where play means they clicked a link and checked it out not because of what the game was but solely because of how it was made.
In terms of concurrent players that game wouldn’t have cracked the Top 5,000 on Steam.
My metric for AI games is “number of players who spent more than 15 minutes playing”. I’m not aware of any vibeslop that has achieved 1 such player.
Now obviously LLMs are transformative for game dev. But “hyper custom worlds you can drop into” shows an extreme ignorance of what players want imho.
> @elonmusk 13h
> Yah
> 158 replies, 74 reposts, 1400 likes
Thanks Elon, you goofy fuck
Maybe nozzlegear wanted to suggest some importance on Musk remaining a bet-ter on the general tech, regardless of the competition?
> "Yah": slang spelling of the word "yeah" (which of course can also be used ironically)
> Merriam-Webster: "Yah": used to express disgust, contempt, defiance, or derision; probably imitative of the sound of retching
Testing, assessing, tasting...
We also do it when we build other things - this is just a different scale.
You cannot come and place your personal positions as assumptions. To me, there is absolutely no theft. And we cannot play a game of "Yes!"//"No!" here.
By the way: are we having a surge of this?
We do have a surge of pro-AI sealions, yes. Any objection is countered with one or more three word questions.
Very devoid of intelligence note.
> Any objection
Objections are arguments. That post did not start an argument - it was as ideological as the Brigades. Devoid of what we want to have here (I believe).
I have stated and do state: what is published is assumed as read (only, probably not read for lack of resources). If it is in the libraries, it is there to be read.
No bias. No speculation. Just facts!
LLMs should be tested in the same way people should be tested for a job interview (but often aren’t) - with tasks RELEVANT to usage.
So you don’t just randomly pick some random thing to make the LLM randomly do (like many job interviewers do).
You start with clear statements about real world usage scenarios. THEN you come up with tests that give insight to how well the LLM/hob seeker gets the job done.
Please, stop coming up with random tests like it’s Microsoft in 1990 and you’re asking job seekers how the would move Mount Fuji, as a way of assessing their programming skills.
No stupid irrelevant pelicans on bicycles and no stupid renderings of Lord Of The Rings. Unless those are relevant use cases.
Any test that anyone comes up with must clearly state the context and how the outcome is measured.
But this demonstrates we're a couple orders of magnitude away from generating 1:1 hyper-personalized entertainment and media for individuals, rather than the masses.
On top of that, a decent chunk of the joy of entertainment is the social aspect.
Ultimately, most people don't have ideas for the kinds of personalized entertainment they want, and they don't want to be in charge of content production (even if you have an LLM do most of the work). I don't doubt that there are niches for it, especially stuff like porn, and I'm sure that pros (game studios, film studios) will leverage AI more and more, but I suspect that most of us will just want to sit on the couch, watch Spiderman XVIII, and then be able to talk about that shared Spiderman XVIII experience with all our friends.
LLMs produce an extremely small range of artistic expression, and its kind of ironic, that while it certainly looks functionally good. It immediately becomes noise because everything looks identical.
So companies are going to have to hire creatives again, in order to stand out, and be "creative"
And this personal media thing, yeah maybe for terminally online persons that are REALLY into a sub-genre but otherwise, it's too much effort, I agree. Until we get machines that can read our (subconscious) mind, that will not exist.
They were visually bad but honest and sometimes soulful or playful.
The AI slop that's everywhere now looks superficially more professional but it's very busy, samey, unnatural and it really turns me off.
Not even close. Current AIs have very poor spatial awareness, they can generate some kind of scene but they can't tell you where objects are in the scene, nor can they move objects into different places. They're very useful but it's difficult to be very creative with them because they can't update an image to make it more aligned with your vision for what should be in the image.
1:1 AI entertainment probably won't be a cold start experience.
Much like the "Choose Your Own Adventure" books of the 80's, consumers choose a baseline template, customize the characters, and interact at various points within the plot.
No, we don't. It's quite good, but it's nowhere near perfect. For example, you can see: https://genai-showdown.specr.net/ or others that highlight how far from perfection we still are.
> and it hasn't really changed the nature of human expression.
It sure has changed our discourse. Look at hackernews, where people just can't help themselves to engage in ragebait and flamewar comment threads about llm generated accusations. We've got half the people convinced that looking at ai content is like consuming food.
I'd venture a guess that it's too early to determine how this technology will change expression. A good analog would be photography, which caused a similar meltdown in the arts at the time.
> I don't see my friends getting wildly creative.
Maybe you don't have creative (enough) friends.
The fans were real heart broken about this but I think you're right on where this is going. We're not going to be seeing the dominance of centrally produced content like this for much longer, like sure, I think there will be big blockbusters will stick around, but I think the day is coming where media becomes a choose your own adventure sort of scenario.
It'll be interesting to see where this scales to. there will definitely be some amazing solo projects but we'll also see the like 4 player co-op version of productions and then the larger mine-craft server 'Minas Tirith' scale ambitious projects that involve a few dozen people. And of course passive consumers will remain a thing, or people who just provide some suggestions or nudges for what they'd like to see others make.
But I don't think it'll be dominated by big companies like Disney, Netflix, Amazon or Paramount.
I'm sure a lot of them will suck but it'll be neat to see the inevitable Seinfield - Star Trek Voyager cross over episodes.
Elaine and B'Elanna Torres feud after a transporter accident leaves the crew stranded the delta quadrant. Jerry attempts to date 7of9 but is rebuffed as she finds Kramer's quirky bluntness more relatable. George panics after someone compares him to Neelix.