the build blog

a making-of ·

Paws in Peril in 3D: a realistic world in the browser, one friend at a time

The obvious way to build a game world is to rough in everything, then polish. We tried that. It gave us 44 weak characters and a world nobody wanted to walk around in. What worked was the opposite: one thing at a time, done properly, then the next.

Version 0 is live: Lola the Bear in a forest clearing, third person, on a laptop or a phone. Two days, me and a team of coding agents, about $1 in model spend. Every section below opens and closes. Jump to the pipeline or the exact prompts.

1Where we started vs. where we ended

Five versions of the same world. The jump from the second to the fourth is the whole post.

All 44 Paws in Peril characters as blocky coloured stick rigs standing in a ring on a flat green plane under a pink sky.
1. The stick-rig ring. Every character at once, as boxes. Proved the plumbing, nothing else.
The 44 characters as procedural toon figures with name tags, crowded together on a flat green field, with buttons for expressions above them.
2. The first procedural shapes. 44 toon figures with faces and moods. In my words: too many characters, too weakly made.
The pilot 3D model of Lola from the front and three-quarter view: a dark, ragged mesh with the sword fused across her body and a hole in her forehead.
3. Lola's pilot model. Straight from the card art, with the pose and the sword baked in. A hole in her head.
Lola's final model from the front and three-quarter view in a clean A-pose: pink, purple and white fur, brown leather armour with steel buckles, a torn berry skirt.
4. Lola from the A-pose redraw. Same card, one cheap redraw first. Clean limbs, real materials, ready to rig.
Lola the Bear waving at the camera on a dirt path in a sunlit meadow, her sword on her back, the log cabin and a forest behind her.
5. Lola, playable in the clearing. Run, jump, a four-hit combo, sit, wave, cheer, dance. Play it. This is the one.

2The biggest lessons

Seven of them, each a before and after. Open the ones you care about.

Breadth first failed. Sparse to dense worked. cuts 1 and 2

The first plan was the "sensible" one: get all 44 characters in, all the systems running, then raise quality across the board. Two rounds later we had 44 characters that all looked like placeholders, and no single thing worth showing anyone.

So we flipped it. Think of a painter: one element at a time, each one great, sparse to dense. One character. Then one tree. Then grass, ground, a cabin, a sky. Each gets its own test stage, its own screenshots, and a separate art-director agent that scores it out of 10 and won't let it through below 8.

Lola standing in tall grass in a sunlit forest glade, rendered in the browser with soft shadows, fur and a blurred background of trees.
The bar. Lola alone, rendered in the browser. One of these beats 44 of the others.
The A-pose redraw: $0.04 that fixed rigging 4.6% torn → 0.14%

Card art is posed. Lola's portrait has her clutching a sword across her body. Turn that into 3D and the arm comes out welded to the skirt, so the first time she waved, 4.6% of her mesh tore.

The fix costs four cents. Before making the model, ask an image model to redraw the card standing in a neutral A-pose: arms 35° out, legs apart, flat light, plain grey. A stick-figure skeleton goes in with the portrait, so the pose is exact. Two takes per character: empty hands (the cleanest rig) and the weapon held out to the side.

Lola's card portrait next to two A-pose redraws on grey: one with empty hands, one holding her sword out to the side.
The card, take 1 (empty hands), take 2 (sword held clear).
Two strips of Lola waving: on top the pilot model, dark and tearing at the arm; below the model from the A-pose redraw, clean through every frame.
Before: the pilot's wave tears. After: 0.14%, nothing you can see.
The best local image-to-3D on an 8 GB laptop Pixal3D, not TRELLIS.2

I wanted models made on my own machine: a laptop with an 8 GB RTX 3070 Ti. We ran the two strongest open image-to-3D models on the same picture.

  • Pixal3D (8-bit weights) worked. About 16–20 minutes and 6.1 GB of VRAM per try. Three seeds, pick one by eye.
  • TRELLIS.2 came out as a flat slab in our setup, every time. We stopped spending time on it.
Two rows from the same card: Pixal3D's model of Lola from front, three-quarter and back on top, and below it TRELLIS.2's output, a flat smeared square.
Same card, same day. Pixal3D (top) vs. TRELLIS.2 (bottom).

Pick the seed for the face first, then the colours, then clean limbs, then the back. A model that doesn't read as the character on the card is a reroll, not a fix.

Rigging a generated mesh the double skin, then the weights

Generated meshes look solid. They aren't. Lola's came out as two skins about 3 mm apart, full of pinholes, with a hollow head and a hole in the forehead. Weight-paint that and the inner skin pokes through the outer one the moment she moves.

  • Make it one solid. Voxelise, seal, then keep only what's visible from 26 directions around her (a flood fill leaks into the hollows) and rebuild the surface. The invented red-panda tail got cut here too.
  • Two sizes: 60k triangles for close-ups, 12k for the game, with the colours baked straight onto clean UVs.
  • Pick joints by eye. Pose detectors missed her knees and eyes; chibi proportions throw them. Fifteen minutes on a gridded front and side view beats it.
  • Voxel-heat weights. Bone heat solved on an 8 mm voxel copy of her, smoothed, head rigid except the ears. 36 bones, including a jaw, eyes, ears and an eight-bone skirt.
  • Then measure. Count torn edges on a walk and a wave. Lola: 0.05% and 0.14%.
Lola's cleaned model from the front, three-quarter and back, and a close-up of the back of her head showing sculpted white and pink fur.
One solid skin, front to back. The close-up is the back of her head, which the game camera stares at all day.
Free motion capture on a chibi bear CC0, $0

No hand-keyed animation. The moves come from the Quaternius animation library, CC0, packaged by the Mesh2Motion project: 178 clips on one human rig, about 13 MB. Free to use, no account.

The catch: those clips are for an adult human. Lola is 1.3 m tall and mostly head. Her hands top out at 1.14 m and her head at 1.27 m, so anything overhead goes through her face. So every clip gets a chibi pass:

  • Feet locked on contact, so she doesn't skate.
  • Hands kept outside a sphere around the head. An overhead cheer becomes a beside-the-head cheer.
  • Arms held 15–20° off the body to clear the belly and the skirt.
  • Exaggerated 12% (emotes 20%). Big heads need big moves to read.
Lola from behind mid-combo on a forest path, sword in hand, pines and beeches all around in late sun.
Mid-combo on the forest path. 19 moves from the library, draw and sheathe made by hand.

Honest status: idle, the wave and the dance are clean. The attacks and sitting still stretch where her arm meets the skirt. Zero torn edges on every move is the bar for v1.

The forest kit: what made it look real eight review rounds

Every tree, blade of grass and log in the clearing is built in code. A few things did most of the work:

  • Light first. A real photographed sky (a CC0 golden-hour HDRI) lights everything, with its sun pulled out into one light that casts the shadows. Nothing looks real under bad light.
  • Dark insides in the trees. Leaves near a crown's core get shaded; the tips get sun. This one change did the most for "that's a real tree."
  • Scanned textures, never painted. Bark, ground, stone and planks are CC0 photo scans.
  • Grass by the blade. Up to 1,800 blades per square metre, each its own height, lean and colour, a few dry. Real modelled daisies, buttercups and clover.
  • The cabin from parts. Notched logs, every cedar shake placed with its own tone and moss, real window glass over a little room inside, grime up the footing.
A stony dirt path winds through a tall meadow to a small log cabin with a cedar roof and a porch, under broadleaf trees and pines.
The clearing from the path.
The view from the cabin porch past a log post and rail: a picket-fenced garden of red clover and daisies, the path, the meadow and the forest beyond.
From the porch. The garden, the path, about 1,400 trees.
Will it run, and what does it cost? measured, not guessed

What costs frame rate is pixels and overdraw, not triangles. Leaf cards, grass, anti-aliasing, bloom and shadows. Lola plus eight more characters added under 2 ms. So the settings tiers cut foliage and effects, not detail on the characters, and a watchdog steps the quality down if the frame rate sags.

DeviceSettingFrame rate
RTX 3070 Ti laptop, 1080phigh165 fps (the screen's cap)
Radeon 680M built-in graphics, 1080pmedium55 fps
Phoneslow, automaticabout 60 on a recent iPhone (estimate)
The game on a laptop: Lola from behind on the path into the woods, her sword on her back, a panel of keyboard controls in the corner.
Laptop. WASD, Shift to sprint, Space to jump, click to attack.
The game on a phone held upright: Lola walking down the forest path, with a Menu button and round Attack, Jump, Draw, Sit and Emote buttons.
Phone. A floating stick and five buttons, on the low setting.

Hosting: $0, at any number of visitors. It's single player, so the whole world is static files on Cloudflare. The whole clearing is a 14 MB download, Lola 1.9 MB of it. Serving the same files from a regular server would cost about $240–1,200 per million visitors.

Multiplayer is the only thing with a meter. A relay on Cloudflare, per month:

Players at onceHow they playPer month
10exploring, 4 h a day$0 (free plan)
100exploring, 4 h a day$6.47
100fighting, around the clock$43.73
1,000exploring, 4 h a day$21

Making a character on the laptop: 1–2 cents of electricity. The whole cast of 44, three tries each: about $3.

3The pipeline

Five steps, A to E one character or one element at a time
  1. A. The redraw. The card art plus a stick-figure skeleton in, an A-pose reference out. Reject any take that changes the design.
    Approve
    One take per character.
    Cost
    $0.04 a picture. All of ours: $1.
  2. B. The model. Cut out the background, run Pixal3D locally on three seeds, compare them on one sheet beside the card.
    Approve
    The seed that reads as the card.
    Cost
    About an hour of laptop, a few cents.
  3. C. Clean and rig. Make it one solid, two sizes, bake the colours, pick joints by eye, voxel-heat weights, then the tearing check on a walk and a wave.
    Approve
    No tearing you can see up close.
    Cost
    $0. About an hour of judgement.
  4. D. Motion and personality. Retarget the CC0 clips, the chibi pass, then a personality profile for how she should move: Lola is heavy (a waddle, a wide stance), brave (chest up) and cozy (she sits when she's bored). V0 has the clips; the profile comes next.
    Approve
    A strip and a looping GIF per move.
    Cost
    $0.
  5. E. Into the world. Each piece gets its own test stage in the browser, automatic screenshots on desktop and phone, an art-director agent to 8/10, then my yes. Then it goes into the clearing and the frame rate gets checked again.
    Approve
    The best two to four renders, on one sheet.
    Cost
    $0 to host, forever.

The scripts and the method are saved as agent skills in the repo, so the next friend (Casper, Speeder, Gropulous) goes through the same five steps.

4The exact prompts

What we sent and the settings we ran copy-paste ready

1. The A-pose redraw (step A). Grok Imagine 2.0, 2:3, 2k. Pictures in: the card portrait, then a stick-figure skeleton in OpenPose colours.

Redraw the character from the first image as a full-body 3D-model reference sheet image, standing upright in a neutral A-pose exactly matching the coloured stick-figure skeleton in the second image: arms held straight out and down about 35 degrees away from the body so they do not touch the torso, legs straight and slightly apart, facing the camera directly, symmetrical. Hands open and empty: remove every weapon and held item (no sword, gun, staff, orb or tool). Nothing in front of the body. Keep exactly the character's design, species, face, eyes, fur or skin, markings, colours, clothing, armour and accessories from the first image. Even, soft, flat studio lighting with no cast shadows and no rim light, plain light grey background, the whole figure visible head to toe with a small margin, painted chibi game-art style matching the first image. No text, no skeleton lines, no other characters.

2. The generation settings (step B). Pixal3D in a private ComfyUI on the laptop.

input     the take-1 redraw, background cut with birefnet-general, 3% margin
model     Pixal3D, 8-bit weights
shape     upsample 1536, remesh 768
seeds     1, 2, 3   (16-20 min each, 6.1 GB VRAM peak)
then      solidify, hero 60k tris + 2048 maps, game 12k tris + 1024 maps

3. One retargeted clip (step D). Blender, run headless.

move      wave, from the library's "Greeting" (CC0)
bones     pelvis→hips, spine_03→chest, upperarm→upperArm, calf→lowerLeg, foot→foot (fingers dropped)
rest      the source arms turned 30° down, from its T-pose to her A-pose
transfer  world-space rotation changes, parents first
root      scaled by her leg length (0.35)
chibi     feet locked, hands outside a 0.28 m head sphere, arms 15-20° off the body, ×1.2
bake      30 fps, trimmed to 3.2 s, then count the torn edges

4. The art director (step E). A separate agent that didn't build the thing, given the screenshots and the references.

You are the art director for a realistic, storybook-warm 3D world rendered in a browser. Score each image 1-10 against the references and against the bar "looks like a frame from a high-end real-time game or a good offline render; nothing reads as procedural, placeholder or plastic". 8 means Christian would be proud to show it. For each image: the score, the three changes that would raise it most, and anything broken (z-fighting, popping, tiling, black artefacts, wrong scale). Then one overall score. Be specific and unsentimental.

5What it cost, and what's next

About $1 in model spend: 25 A-pose redraws at four cents. Everything else ran on my laptop or came free under CC0. Hosting is $0.

Next up: zero tearing on every move, then one more friend at a time, then more houses and plants, filling the world outward from the clearing. Go say hi to Lola.