We are vibe coding my game. It did it entirely in 3d in 5.5 and we asked it to play test it for us. We actually watched it play the game and it made observations on what needed fixing. It's a crazy frontier.
It really is amazing, and also amazingly frustrating when it can knock out all this functionality no problem, and then you hit something adjacent that it just can't do. Very possibly/probably a skill issue on my part, but with whatever my skill is or isn't, it can do some amazing stuff with ease and then hit a wall. I may not understand the technical bits behind how 3d actually happens well enough to communicate it and have it tested appropriately in spots where the ai can't figure it out itself.
My last post probably undersold the 3d ability, but that's where I keep getting it stuck or at least bogged down. I can only assume that It innately understands typical coding issues from basically being trained on all the public coding knowledge in the world, but for 3d, there's little/no innate understanding of what the typical problems and solutions are, so every little thing has to be derived ground-up and managed in context.
To illustrate what I mean, here's a dumb summary of my dumb 3d adventures (2d seems to fare much better, but still a ton of issues with consistency), skip unless you're also dumb and want to have ai make a game: (and feel free to pile on the dumbness if you can also tell me something that might decrease it lol)
(I should note that not all of these steps have been retried on 5.5 yet.)
Throw some 3d stuff in a scene and have it move around? no problem. Now build the models you're moving around? Eh, ok kindof, for general llms at least. The dedicated model makers seem reasonable.
Play a premade animation or load a premade pose? No problem. Humanoid animation or pose from scratch? Eh, it can technically move the bits, but in a way that looks good/natural? No.
So the control is there but not the judgement/validation that what it's doing is in any way reasonable, so practically speaking, here's what I run into.
Ok, here's a rigged character and a model of a chair, put the model into a natural looking seated pose. It kind of works at a placeholder or super casual, low-poly level.
Ok, here's a library of poses, can you analyze and rank these for which ones might be a likely candidate for sitting in this chair? Works a little, but super scattershot.
Here's a set of rules that will help you differentiate sitting poses. Also, here's a big list of confirmed sitting poses, can you analyze them for similarities and add those to your rules list, then rerank the library. Big improvement, but still fails a lot at selection.
Now use the pose on the chair. NO NOT LIKE THAT! Can you identify the chair back since that should imply orientation? "Absolutely" Fucking liar. Ok, we've added zones to the chair that identify seating area, back, and collision. Orient the character opposite the seat back, and reject any pose that causes a bone to collide with a collision area. Helps, but still a ton of fails.
Ok, bones are hard, let's evaluate against collider capsules. Better, but still fails very often.
At this point, this seems like it should be completely mechanical/deterministic to to me. It kind of works but still fails so often and badly that I'm just confused.
Ok, but a high failure rate is no problem if the ai knows when it failed. But it does not. Ok, if you can't judge from the 3d data, take screenshots from different angles and judge them. Ok you suck at that, so pass them to a dedicated visual model for evaluation. They suck at that particular thing too. They can tell you what's there, but not if what's there is being used appropriately.
Here are mocap libraries that analyze/extract bone positions from real photos/videos. Let's try that. This one seems like it should reasonably work since there are products based on them that seem to work, but I've gotten pretty terrible results. This one is confusing, feels like I'm missing an essential piece somehow.
Ok, you can get some of the way there, but still lots of collisions. If a collision is minor nudge the bones until they stop colliding. and feet shouldn't float, so nudge the bones until they're in contact. Here are some rules/criteria. This actually kind of works, but it took an unspeakable number of iterations. and is still basically worthless if the changes need to cascade though multiple bones.
Ok it kind of works. let's try to generalize to different furniture/positions.
The latest adventure is cloth sim, which is making the above rant seem like an amazing success in comparison!