I built a simple Rust game (not my first one) and wanted to see whether Laya could actually play it. I’d seen Jev playing Minecraft, and small-model demos in games like Snake. Time to test the hype in a game of my own.
The game gave me a fairly small problem to start with. Enter a room, kill the enemies, find the next room. Keep doing that until the whole floor is clear. No inventory management or elaborate quest to explain away a failure.
The game is a Binding-of-Isaac-style prototype. Doors seal during combat and reopen when the enemies die. Chasers move toward the player and hurt on contact. Winning means clearing every combat room, not surviving until the clock runs out.
Laya receives structured room data, not screenshots. The game pauses during inference, so a slow model call does not leave the player standing there while an enemy kills it.
What one Laya call looks like
Text and a list of answers go in. One score per answer comes out. This is a real call from a fight later in the post.
Code sends
State
Goal: clear all combat rooms. Stay alive. … Health: 3. Phase: combat. Clearance now/goal, tiles: enemy 11.46/4; wall 0.1/2. Clearance goal unmet: yes. Chaser 0: health 2, distance 11.46, aligned false. Chaser 1: health 2, distance 22.03, aligned false. Chaser 2: health 2, distance 15.56, aligned false. …
Question
Create space when clearance is inadequate; otherwise focus a nearby weakened chaser.
Answers
focus_0: Retain 0; position/fire to defeatfocus_1: Retain 1; position/fire to defeatfocus_2: Retain 2; position/fire to defeatcreate_space: Seek clearance; nearest fire
Laya returns
focus_09.7%focus_18.9%focus_212.4%create_spacepicked69.0%
Code takes the top answer and acts on it. Here, Rust starts moving the player away from the wall.
There was one neural model in this experiment. I ran the pinned English Laya checkpoint locally through MLX on Apple Silicon. Jev was part of the inspiration, not another player I tested. There was no game-specific training or fine-tuning. All the changes below are changes to prompts, action selection or the code around the same weights.
Auto-aim wasn’t enough
Laya died with auto-aim turned on.
In one run, it had one health left and was pressed against the east wall. A chaser approached diagonally. Laya kept choosing to stay and fire west. It chose stay again. The chaser killed it on the next tick.
Rust already aimed the gun. It pointed at the nearest enemy in a cardinal direction, west here. Laya still chose where to stand. Pointing west did not make that spot safe.
To know whether Laya was any good, I wrote two bots to beat. One followed deterministic Rust rules. The other picked legal actions from a fixed seeded random stream. Neither used a neural model.
All three got the same Rust aiming, the same doorway steering and the same short-action menu. Each command ran for at most ten ticks at 60 Hz, about a sixth of a simulated second. Then came another choice.
Where Laya sits: short actions
Each layer runs inside the one above it. Laya picks the route and the move. Rust only aims and steers.
Objective whole floor
Clear every combat room. The game sets this.
Goal none held
Nothing keeps a target or plan between calls. Each call starts again from the room state.
Route next doorway
Pick an open doorway or wait. The next call can pick a different doorway 10 ticks later.
Move ≤10 ticks, about 1/6 s
In combat, stay or move in one of eight directions, always firing.
Tick every 1/60 s
Aim, steer onto the chosen doorway, run physics.
We ran them through the same ten maps, with the same physics and a 120-second simulated-time budget. Waiting for Laya did not use that time. The Rust bot cleared every floor without damage. Laya and random cleared none.
Laya against two plain bots
The first ten-map test, before skills.
| Bot | Wins | Deaths | Out of time |
|---|---|---|---|
| Rust rules | 10/10 | 0 | 0 |
| Laya | 0/10 | 3 | 7 |
| Random short actions | 0/10 | 0 | 10 |
The Rust bot got there with fixed rules. It chose the least-visited destination and scored combat moves for space from chasers and walls, plus firing alignment. That was static geometry, not the physics-prediction controller built later.
In a different recorded fight near the south wall, Rust moved northeast while firing north. Ten ticks later, it had almost lined up with the chaser. Then it stayed and fired. Staying was useful in that position. It was lethal in Laya’s east-wall fight.
Random found a different way to fail. In one run, it kept changing destinations until its time ran out. It had not killed an enemy or cleared a room.
Laya cleared three isolated combat tests without damage. Rust did too. Random cleared them with damage. Laya could also cross a doorway or survive a short test. But on the full floor, it still died or ran out of time.
Maybe the prompt is bad
Laya was still dancing around the starting room. I suspected the prompt.
That was a reasonable place to look. Laya got the same dump every time: positions, enemies, exits and its last few actions. Much of it didn’t matter. In a fight, it still read about sealed doors it couldn’t use. While exploring, it saw how often it had visited each neighbouring room, but not whether a cleared room led anywhere new.
So I stopped sending a fixed dump. Before every call, the game code looked at what was happening and built the prompt from that. A small behavior tree does the looking, built with Bonsai, a Rust library. It checks whether the player is moving between rooms, fighting, or exploring, then picks what goes into the prompt. A fight gets enemies, walls and shots. Exploring gets the open doors and a map of the rooms the player has already seen, so Laya can tell a dead end from a door that still leads somewhere new.
That map holds only what the player has seen. It is not the hidden floor layout, and nothing was trained.
There was a second problem, and it had nothing to do with the prompt. Laya picked a door, the bot started walking, and a sixth of a second later Laya could pick a different door. At the start of one run, it chose south, then west, then north. Each new pick undid the last few steps. Even a good pick could look wrong on screen: to reach a west door, the bot first has to walk north to line up with it.
So in the next version, Rust kept Laya’s door until the bot got there, or until the trip went wrong, like getting stuck or taking too long. Rust already knew how to line up with a door. Now it got to finish the walk. This changed how the bot moved, not just what Laya read.
Full floors · 10 per version
Describe the situation. Keep the destination.
Do a better prompt and kept doors clear more floors?
-
A · Original prompt
0/10 wins 3 deaths 7 timeouts
-
B · Phase prompt + map
1/10 wins 3 deaths 6 timeouts
-
C · B + keep the door
1/10 wins 4 deaths 5 timeouts
Limit. B and C win only the same single map. Neither gets close to 8 wins out of 10.
What changed, with prompt excerpts
Each version: the same ten maps, 120 simulated seconds, same model, physics, allowed moves and auto-aim. C also changes how the bot walks to a door.
A · Original prompt
Every position, every door, visit counts and recent moves. Pick again every sixth of a second.
Clear every combat room. Explore open doorways. Stay alive.
Axes: east +x, south +y. Chasers hurt on contact.
tick=0 room=0 bounds=-16,16
player position=0,0 velocity=0,0 health=3
exit North position=0,-16 sealed=false visits=0 B · Phase prompt + map
Only what matters right now: fighting, exploring or changing rooms. A remembered map shows which doors are unexplored, dead ends or a way through.
Choose the action that advances clearing ALL remaining combat rooms. Explore unmapped doorways or reach unexplored branches. Repeating cleared dead ends or waiting cannot win.
North option text:
Head North to explore an unmapped doorway; short step, then choose again; no fire C · B + keep the door
Keep the chosen door until the bot arrives or the trip fails. Fights still use short moves.
North option text:
Head North to explore an unmapped doorway; retain target until arrival or interruption; no fire Together, the new prompt and the remembered map took wins from zero to one. Keeping the door won that same map again, and no other. Keeping a door only helps if it’s the right door, and neither change touched the fights.
Rewording the question didn’t help either. Asking straight out which door led somewhere new made Laya more sure about the useful doors, but it didn’t pick them any more often. That test, and two more, are in “The other prompt experiments” below.
Escaping one failure finds another
The fights were the bigger problem. My first guess was the words. Maybe “move north” was hard to use when the gun fires west, so I tried names like “sidestep right relative to aim”. The moves behind the names stayed the same. New descriptions fixed some bad choices and broke some good ones. New names were worse.
So maybe Laya couldn’t turn coordinates into facts. Rust could. The next prompts said things like “east is blocked” and “northeast slides along the wall,” plus whether a chaser stood in the firing line. They described the room right now. They didn’t say which move was safe.
Saved fight moments · 47 per version
Spell out the walls and the aim
Do wall and firing-line facts prevent damage?
Saved moments replayed one at a time, not full games.
-
A · Original
38/47 took no damage
- Fixes
- 0
- New damage
- 0
-
B · Blocked + sliding moves
42/47 took no damage
- Fixes
- 6
- New damage
- 2
-
C · Firing line
37/47 took no damage
- Fixes
- 2
- New damage
- 3
-
D · Both
39/47 took no damage
- Fixes
- 2
- New damage
- 1
Fixes count the 9 moments where the original prompt took damage. New damage counts the 38 where it didn't.
Limit. The best version still made 2 new damaging choices. 4 of its 6 fixes came where no move was blocked, so the wall facts may not be why it improved.
What changed, with prompt excerpts
The same 47 saved fight moments and compass moves. Only facts about right now are added, never which move would be safe.
A · Original
Positions, wall distances, motion, health and cooldowns.
Walls: north 11.83, south 20.17, east 0.35, west 31.65.
Chaser 2: 0.27 south, 0.95 west; motion 1.49 north, 5.3 east. Health 2. B · Blocked + sliding moves
Add which moves are blocked or slide along a wall right now.
Wall contact now: blocked moves east; sliding moves ne, se. Blocked moves can still fire. C · Firing line
Add whether a chaser is in the current firing line. Not a prediction of a hit.
Current West firing lane: nearest chaser 2 is inside it. Current geometry, not a hit prediction. D · Both
B and C together.
Wall contact now: blocked moves east; sliding moves ne, se. Blocked moves can still fire.
Current West firing lane: nearest chaser 2 is inside it. Current geometry, not a hit prediction. The best version still made two new damaging choices. Here is one. At the east wall, the original prompt led Laya to pick north, and it took no damage. With the wall facts added, it picked northeast and got hit two ticks later.
A northeast command at an east wall
Recorded state from a development run. This is not current gameplay footage.
- P · Player radius
- 0.35 tiles
- C · Chaser radius
- 0.55 tiles
- Chaser relative to player
- 0.95 west
0.27 south - Player center to east wall
- 0.35 tiles
The wall blocks the east component. Northeast slides north more slowly than a pure North command. Arrows show directions, not future paths.
North: no damage over 10 ticks
Northeast: hit after 2 ticks
Exact prompt excerpts
Original prompt
Walls: north 11.83, south 20.17, east 0.35, west 31.65. Chaser 2: 0.27 south, 0.95 west; motion 1.49 north, 5.3 east. Health 2.
Added wall facts
Wall contact now: blocked moves east; sliding moves ne, se. Blocked moves can still fire.
A diagonal move splits its speed between two directions. The wall blocks the east half, so the player slides north slowly, and the chaser catches it. The prompt even said northeast would slide. Describing the room is not the same as knowing what each move does over the next few ticks.
On full floors, the wall facts stopped the deaths, but the bot ran out of time on nine floors instead. It was going in circles. Once it walked back into a cleared room it had already visited 153 times, while two other doors still led somewhere new. A short history of recent door trips didn’t stop it. The bot walked there just fine. Moving wasn’t the problem. Choosing where to go was.
Laya kept walking back into the same dead end because that door was always its top answer. But Laya doesn’t give just one answer. It scores every door. So what if the bot rolled dice weighted by those scores? A door with a lower score would still get picked now and then, and that might break the loop. I changed only that, and only while exploring. Fights still took Laya’s top answer.
Full floors · 10 per version
Roll dice instead of taking the top answer
Can Laya's own scores break the loops?
-
A · Top answer
1/10 wins 0 deaths 9 timeouts
Totals across all 10 floors
- Rooms reached
- 20
- Rooms cleared
- 20
- Revisits
- 2,762
-
B · Weighted dice
1/10 wins 9 deaths 0 timeouts
Totals across all 10 floors
- Rooms reached
- 31
- Rooms cleared
- 22
- Revisits
- 128
Early-death warning. Revisits and model calls drop partly because 9 runs die early. One dice sequence per map. The only win is still on the same map.
What changed, with prompt excerpts
Same ten maps, doors kept until arrival, wall facts on, 120-second limit. While exploring, every allowed choice, wait included, goes into a weighted dice roll. Fights still take the top answer.
A · Top answer
Always do Laya's top exploring choice.
B · Weighted dice
One weighted dice roll per new exploring choice. Same prompt, no rerolls, no anti-loop rule.
The nine timeouts became nine deaths. The number of wins did not change.
The dice worked, in a way. The bot stopped circling cleared rooms and walked into new fights. That was the problem: walking in circles had been keeping it out of fights it couldn’t win. In one of those deaths, the player had one health and a chaser sat just over a tile to the south. Laya chose to move east, even though the prompt said to avoid contact. A twentieth of a second later, the run was over. Room revisits fell from 2,762 to 128, but dead bots stop revisiting rooms, so that isn’t progress.
The bot now failed in a new way, but it still failed. Better inputs, clearer wording and weighted dice had not taught short-action Laya to survive a fight.
The other prompt experiments
Three more tests from the same stretch. None of them changed the story, so they live here.
Saved exploring moments · 32 per version
Four ways to ask for an unexplored door
Does a narrower question or plainer wording fix the choice?
Saved moments replayed one at a time, not full games.
-
A · Original
27/32 headed somewhere new
Average score on useful doors 76.65%
-
B · Narrower question
27/32 headed somewhere new
Average score on useful doors 81.07%
-
C · Plainer door text
27/32 headed somewhere new
Average score on useful doors 76.62%
-
D · Plainer room text
27/32 headed somewhere new
Average score on useful doors 76.35%
Limit. All four fail on the same five moments. The narrower question raised Laya's scores for useful doors, but it didn't pick them any more often.
What changed, with prompt excerpts
The same 32 saved exploring moments, with the same doors in the same order. Only the wording changes. Each version answered three times per moment.
A · Original
The original question and dead-end wording.
Question:
Choose the action that advances clearing ALL remaining combat rooms. Explore unmapped doorways or reach unexplored branches. Repeating cleared dead ends or waiting cannot win.
West option text:
revisit a cleared room with no known unexplored branch
Room description:
No known unexplored branch without returning here. B · Narrower question
Change only the question. Everything else stays.
Which doorway offers an unknown connection or a known route to an unexplored branch? C · Plainer door text
Rewrite how a dead-end door is described. Nothing else changes.
Before:
revisit a cleared room with no known unexplored branch
After:
revisit a cleared room; further exploration returns here D · Plainer room text
Rewrite the dead-end sentence in the room description. Not combined with C.
Before:
No known unexplored branch without returning here.
After:
Further exploration requires returning here. Saved fight moments · 47 per version
Rename the moves, keep the physics
Are moves named relative to the gun easier to pick than compass directions?
Saved moments replayed one at a time, not full games.
-
A · Compass moves
38/47 took no damage
- Fixes
- 0
- New damage
- 0
-
B · Relative descriptions
41/47 took no damage
- Fixes
- 6
- New damage
- 3
-
C · Relative names
27/47 took no damage
- Fixes
- 0
- New damage
- 11
-
D · Both
32/47 took no damage
- Fixes
- 1
- New damage
- 7
-
E · D + relative room
32/47 took no damage
- Fixes
- 1
- New damage
- 7
Fixes count the 9 moments where the original prompt took damage. New damage counts the 38 where it didn't.
Limit. New descriptions fixed 6 of the 9 bad moments but broke 3 of the 38 good ones. No version fixed something without breaking something. These are 47 short moments, not 47 games.
What changed, with prompt excerpts
47 saved fight moments where Laya had chosen badly before. Same nine moves, auto-aim and option order. Every moment had at least one safe move.
A · Compass moves
The original move names and descriptions.
Firing West:
north: Move north; fire West B · Relative descriptions
Only the descriptions change. Names and room description stay in compass terms.
north: Sidestep right relative to aim; fire. C · Relative names
Only the names change. Descriptions stay in compass terms.
sidestep_right: Move north; fire West D · Both
Relative names and relative descriptions.
sidestep_right: Sidestep right relative to aim; fire. E · D + relative room
Also describe the room relative to the gun. Same distances.
Walls: right 11.83, left 20.17, behind 0.35, ahead 31.65.
Chaser 2: 0.27 left, 0.95 ahead; motion 1.49 right, 5.3 behind. Health 2. Full floors · 10 per version
Try the fixes on full floors
Do fewer bad moves turn into more wins?
-
A · As before
1/10 wins 4 deaths 5 timeouts
-
B · Wall facts
1/10 wins 0 deaths 9 timeouts
-
C · Recent door trips
0/10 wins 4 deaths 6 timeouts
-
D · Both
0/10 wins 2 deaths 8 timeouts
Limit. Wall facts stop the deaths but leave 9 timeouts. History does not stop the back-and-forth. No version reaches 8 wins out of 10. This run went ahead even though the wall facts had failed the fight test.
What changed, with prompt excerpts
Each version: the same ten maps, 120 simulated seconds, doors kept until arrival, Laya's top answer taken. Same model, physics and allowed moves.
A · As before
Same prompt as the doorway-commitment version.
B · Wall facts
In fights, tell Laya which moves are blocked or slide along a wall. Exploring is unchanged.
C · Recent door trips
While exploring, list the last four door trips and the door the bot came in through.
Prompt excerpt:
Recent room crossings: room 0 -> room 2; room 2 -> room 0; room 0 -> room 2; room 2 -> room 0.
Entered current room 0 through North doorway from room 2.
Laya picked: North, the way back, even though South was unexplored. D · Both
B and C together. Nothing filters or replaces Laya's choice.
What were the Snake and Minecraft bots doing?
I went back to the motivation for the experiment. I’d seen Jev play Minecraft. Why couldn’t Laya consistently complete these much simpler rooms?
So I read the code behind the public Snake and Minecraft bots I could find. I couldn’t find an official TypeSafe one, so these are all community projects. The results below are what their builders reported. I didn’t rerun them.
| Bot | The model picks | Code works out before asking | Code does after |
|---|---|---|---|
| Jev Minecraft, rmalde | One task from a list: walk to a waypoint, mine, craft, loot or fight | A separate language model sets the current goal. The route was mapped out in advance. | A bot library does the walking and mining. A reflex can cancel any task the moment danger shows up. |
| Jev Minecraft, ellistev | Quarter-second steps and 30° turns | Code checks the blocks around the player: clear, blocked or dangerous. Unsafe moves are left out. | An answer is thrown away if it took over 5 seconds, or if the player moved while waiting. |
| Jev Snake, sorrycc | The next square | Deadly moves are removed. Moves into dead ends are marked. | One square per turn. |
| Laya Snake, mizorewww | The next square | A planner labels each option, for example “Safe. Best route to food.” | A safety check swaps an unsafe pick for the safe move Laya scored highest. |
| Laya vs Jev race | Turn left, right or go straight | Where the apple is, relative to the snake’s head | Nothing. Walls wrap around and the snake can cross itself, so it cannot die. |
The Laya Snake port has the strongest result. It ran 8,160 moves with zero deaths, and the safety check overruled Laya only four times. But its planner had already labelled the best move in every list, so that’s a result for planner plus Laya. It also used the multilingual version of Laya, not the English one I used.
The ellistev Minecraft bot is the closest match to my experiment, because its builder tried both ways. When the model picked whole tasks, the bot built a 338-block flag in 590 seconds with 131 decisions. When the model steered every quarter-second step, the same flag took 801 seconds and 1,848 decisions, and the README says the bot can wobble back and forth and get stuck.
A bug report on Laya’s own GitHub shows the same limit from the other side. Each option already said what would happen. Laya still gave right a 53% score, even though the text for right said the snake would hit the wall and the game would end. The maintainer replied that Laya is “a single-pass classifier, not a planner”: it reads the options once and scores them, without thinking ahead. His fix was to remove the deadly moves in code. That report used the multilingual version too.
My short-action bot had none of that help. Nothing removed deadly moves before the question, nothing caught them after, and nothing marked the best option.
I hadn’t specialized Laya for this game, and none of this shows what happened inside the model. But it told me what to try next.
So the next change gave Laya bigger things to choose. Not new names for the same tiny moves, but whole tasks, with Rust doing the walking.
Give it a goal, let Rust drive
So I gave Laya bigger choices. Instead of a direction, it now picked a goal: fight this chaser, get some room, or go explore that unexplored door. Rust kept the goal and did the moving.
That’s a lot more help than aiming the gun and lining up with doors. It changed Laya’s job. The model stayed the same, and nothing was trained.
Fighting a chaser meant Rust kept that target, moved to where it could hit it, fired, and stayed out of reach. Laya no longer had to pick northeast, then stay, then northeast again just to line up one shot.
Getting room had no target. Rust moved away from enemies and walls and kept shooting at whatever was closest.
For both, Rust tried all nine moves every tick and played each one forward about a fifth of a second, using the game’s real movement and combat code. Staying alive came first. After that, a fight move scored on lining up shots on its target from a useful distance, and a room move scored on how much space it bought.
Exploring changed too. Laya now picked an unexplored door anywhere on the map the player had seen, even a few rooms away. Rust found the way there through rooms it already knew and walked the whole trip without asking again.
Where Laya sits: skills
Laya moves up one layer. Rust takes everything below it.
Objective whole floor
Goal up to 1 s in fights, 30 s walking
Fight a chosen chaser, get some room, or explore a chosen door. Rust keeps the goal until it's done, fails or runs out of time.
Route rooms to the door
Find the way through rooms the player has seen. No new question in between.
Move picked again every tick
Play each of nine moves a fifth of a second forward with the real game code. Staying alive comes first.
Tick every 1/60 s
Aim at the chosen chaser, or at the nearest one while getting room.
Rust still didn’t see the hidden map, and it never swapped Laya’s choice for its own. But “move northeast for a moment” had become “kill this chaser without getting hit.” Dodging, positioning and finding the way were now code’s job.
The skill version pushed that idea further. Before every call, code rebuilt the whole question from the game state: what Laya read, what it was asked, and which answers existed at all. An answer only appeared if Rust could carry it out right then. The create_space answer showed up only while the player was too close to a wall or an enemy. A dead chaser’s focus answer disappeared.
The prompt is rebuilt before every call
Four calls in a row from the start of one run, about three seconds of play. Code decides what Laya reads, what it is asked and which answers exist.
Start of the floor
Laya reads
Phase: exploration. Frontier 0 North: route crossings 1. Frontier 0 East: route crossings 1. …
Question
Explore an unresolved frontier with a short known route.
Answers
First fight, pinned to a wall
Laya reads
Phase: combat. Clearance now/goal, tiles: enemy 11.46/4; wall 0.1/2. Clearance goal unmet: yes. Chaser 0: health 2, distance 11.46, aligned false. …
Question
Create space when clearance is inadequate; otherwise focus a nearby weakened chaser.
Answers
A moment later, with room
Laya reads
Phase: combat. Clearance now/goal, tiles: enemy 8.6/4; wall 2.1/2. Clearance goal unmet: no. Chaser 0: health 2, distance 8.6, aligned false. …
Question
Create space when clearance is inadequate; otherwise focus a nearby weakened chaser.
Answers
Chaser 0 is dead
Laya reads
Phase: combat. Chaser 1: health 2, distance 13.69, aligned false. Chaser 2: health 2, distance 9.02, aligned false. …
Question
Create space when clearance is inadequate; otherwise focus a nearby weakened chaser.
Answers
That meant the comparison had to change too. Pitting this bot against the old random bot would give Laya credit for all the new Rust code. So all three bots got exactly the same skills, and only the picking differed.
The fixed-rules bot got new rules for picking goals. It got room when it was too close to something. Otherwise it went for the nearest chaser it could kill in one hit, or just the nearest one. When exploring, it picked the unexplored door with the fewest rooms to cross. The skills did the moving.
Random picked evenly from whatever skills were on offer. It ran five times on each of the ten maps, with different dice each time. That’s fifty runs, but still only ten maps.
Same skills, three ways to pick them
70 wins, none with damage, on 10 development maps.
| Picked by | Wins | Average seconds | Model calls |
|---|---|---|---|
| Laya | 10/10 | 21.56 | 175 |
| Fixed rules | 10/10 | 19.73 | 0 |
| Random | 50/50 | 21.89 | 0 |
Random ran five times on each map with different dice: 50 runs on the same ten maps, not fifty different maps.
- Laya
- Fixed rules
- Random, five runs
Completion time in simulated seconds. Lower is faster.
A dotted line joins Laya and fixed rules on each map. The five diamonds are random's five runs. Rows are spread out only so markers don't overlap; all 70 times use the same 0–45 second scale.
Laya was 8.4% slower than fixed rules, comparing the two map by map. Dividing the table's averages gives a different number.
Against random, Laya was 1.5% faster, but the likely range runs from 8.7% faster to 6.2% slower. That is no real difference.
All 70 raw completion times
Unrounded simulated seconds. Random runs are listed in the order they ran.
Map 1
- Laya:
22.283333333333335 - Fixed rules:
22.333333333333332 - Random run 1:
24.2 - Random run 2:
24.116666666666667 - Random run 3:
23.866666666666667 - Random run 4:
23.45 - Random run 5:
22.566666666666666
Map 2
- Laya:
24.05 - Fixed rules:
19.716666666666665 - Random run 1:
25.15 - Random run 2:
21.05 - Random run 3:
27.666666666666668 - Random run 4:
20.7 - Random run 5:
31.783333333333335
Map 3
- Laya:
21.133333333333333 - Fixed rules:
13.933333333333334 - Random run 1:
16.166666666666668 - Random run 2:
15.083333333333334 - Random run 3:
18.316666666666666 - Random run 4:
15.75 - Random run 5:
19.666666666666668
Map 4
- Laya:
23.733333333333334 - Fixed rules:
24.133333333333333 - Random run 1:
24 - Random run 2:
23.883333333333333 - Random run 3:
23.833333333333332 - Random run 4:
23.333333333333332 - Random run 5:
23.683333333333334
Map 5
- Laya:
14.216666666666667 - Fixed rules:
14.216666666666667 - Random run 1:
14.016666666666667 - Random run 2:
12.6 - Random run 3:
13.633333333333333 - Random run 4:
13.65 - Random run 5:
14.066666666666666
Map 6
- Laya:
16.783333333333335 - Fixed rules:
17.133333333333333 - Random run 1:
17.816666666666666 - Random run 2:
16.95 - Random run 3:
17.433333333333334 - Random run 4:
17.3 - Random run 5:
17.416666666666668
Map 7
- Laya:
38.7 - Fixed rules:
32.75 - Random run 1:
33.46666666666667 - Random run 2:
30.1 - Random run 3:
35.31666666666667 - Random run 4:
40.11666666666667 - Random run 5:
33.983333333333334
Map 8
- Laya:
15.2 - Fixed rules:
15.416666666666666 - Random run 1:
19.166666666666668 - Random run 2:
21 - Random run 3:
16.433333333333334 - Random run 4:
19.516666666666666 - Random run 5:
14.683333333333334
Map 9
- Laya:
17.1 - Fixed rules:
17.65 - Random run 1:
16.916666666666668 - Random run 2:
17.25 - Random run 3:
17.516666666666666 - Random run 4:
17.333333333333332 - Random run 5:
16.683333333333334
Map 10
- Laya:
22.35 - Fixed rules:
20.05 - Random run 1:
25.016666666666666 - Random run 2:
29.983333333333334 - Random run 3:
26.116666666666667 - Random run 4:
32.2 - Random run 5:
28.65
Laya cleared all ten floors. Fixed rules cleared all ten. Random cleared all fifty runs. None of the seventy runs took a single hit.
Random was choosing chasers and doors now, not moves. Earlier, random could turn back from a door almost immediately or step the wrong way next to an enemy. Now a random chaser came with Rust code that approached, shot and dodged. A random door came with a route that got there. And random only ever saw answers that made sense right then, because code had already removed the rest. Each skill had a time limit, but it ran long enough to get real work done before the next choice.
The bot worked. I still hadn’t shown why it needed Laya.
Laya won no more floors, and it was 8.4% slower than fixed rules, comparing each map with itself. It was no faster than random either.
Those are game-time numbers. The game pauses while Laya thinks, so they leave out the waiting. The fixed-rules and random bots never wait at all, because they don’t call a model.
The game now
Current gameplay excerpts
The skill-controlled clip is at the start. Both 22-second clips were captured October 5 and show current controls, not the historical campaign.
Short-action Laya
The game has changed since September, so these clips illustrate the controls rather than reproduce the comparison. The current skill trace won once, in 20.35 simulated seconds with nineteen model calls and no damage.
The historical results also came from development maps that we’d already used to diagnose failures. Held-out worlds remained unused. Reliable on those ten worlds is the result we have, not a claim about every floor the generator can produce.
Takeaways
What the working demos do
- Compute the geometry in code. The Snake demos compute dead ends, reachable cells and the route to food before they ask. TypeSafe’s Jev guidance says it plainly: “Keep the arithmetic in code.” My first bot sent coordinates and wall distances.
- Remove or replace fatal choices. The Jev Snake removes fatal moves before the question. The Laya Snake shields the answer after it. My short-action bot had neither, so the fatal stay at the east wall went straight through.
- Offer actions that already work. In Minecraft, “travel to the waypoint” calls a pathfinder. Our skills became the same kind of action.
- Let code check every tick. The Minecraft breath reflex can cancel an action on any physics tick without asking the model. My game paused during inference, so latency was never the problem. A short-action combat move ran for up to ten ticks. Rust stopped it after damage or a room change, but never predicted danger and picked a safer move. The skill controller does that every tick.
- Send only the facts the question needs. In the Laya-vs-Jev arena’s Flappy Bird game, a state that described both the gap and the bird made Laya answer with the bird’s position, 0/6. Describing only the gap gave 6/6.
- Show how much the code does. The Laya Snake port counts shield interventions. ellistev labels direct and high-level runs separately. Our fixed and random baselines do the same job: they show what the code achieves without the model.
All of these demos gave the model more help than my first bot had. For the skills, I borrowed the high-level Minecraft design. The model picks a goal, and code handles the movement.
What our prompt changes actually did
The version that worked did two things together, so I can’t split the credit:
- Ask for goals, not movement. Bounded skills cleared 10/10 floors with no damage, against 0/10 for short actions.
- Rebuild the whole question from the game state before every call. Code chose what Laya read, what it was asked and which answers existed. Laya never saw an answer Rust couldn’t carry out right then. One of the Minecraft demos does the same: gathering and building offer different actions.
Three changes helped a little. Two of them also broke cases that used to work:
- Phase context and a discovered map. Together they raised floor wins from 0/10 to 1/10. I changed both at once, so I can’t split the credit.
- Blocked moves in words. “East is blocked” and “northeast slides along the wall” took recorded states from 38/47 to 42/47 damage-free, with two new damaging choices. On full floors, deaths fell from 4 to 0, but wins stayed at 1/10.
- Enemy-relative descriptions. Damage-free states went from 38/47 to 41/47, with three new damaging choices.
Three changes did not help. A narrower exploration question raised the probability on useful doorways from 76.65% to 81.07%, but picked correctly in the same 27 of 32 states. Relative action IDs dropped combat from 38/47 to 27/47 damage-free. Recent crossing history took floor wins from 1/10 to 0/10.
These habits kept the results honest:
- Test wording on frozen states with a no-new-damage rule. It exposed the regressions above, which better totals would have hidden.
- Change one thing at a time. IDs, descriptions and option order each moved results. Reversing the exploration option order fixed three states and broke two.
- Don’t ask when there is only one answer. In the first campaign, 3,590 of 6,817 calls offered only a forced wait. The skill version skips those calls.
- Fit the prompt without cutting it. The English checkpoint reads 512 tokens per question. A preflight check rejects a shortened option or a cut instruction, then checks the real tokenized input. Skill prompts used 150 to 249 tokens.
The combat question that worked was short. It asked for one judgment, and every option named something Rust could finish:
Create space when clearance is inadequate; otherwise focus a nearby weakened chaser.
focus_0: Retain 0; position/fire to defeat
create_space: Seek clearance; nearest fire
Protocol and limits
Each comparison kept its settings fixed. All full-floor runs used the same ten development maps and a 120-second limit in game time. Physics ran 60 times a second and paused while Laya answered. Within a comparison, every bot got the same maps, rules, moves and helpers. Later comparisons changed what Laya saw or how moves ran, on purpose, so the results are a series of tests, not one long A/B test.
The first Rust bot used no randomness. When exploring, it went to the least-visited door, breaking ties North, South, East, West. In fights, it tried a two-tile move in each direction and scored it on distance from the nearest chaser, lining up a shot, and distance from the walls. It did not simulate the next few moments like the later skill code did. The first random bot used one fixed sequence of random numbers per map. The skill test gave random five fixed sequences per map, drawn without bias.
The model was the English Laya checkpoint c5d78730f3493e4fe16d61507ef4b78eef7318cf, run through laya-apple 1.5.0 on the Mac’s GPU in 16-bit floats. It reads at most 512 tokens per question. I did not test the multilingual or typed-decisions versions in the game. Every input and action was recorded, so runs can be replayed without asking Laya again. The bot never quietly swapped in a rule when Laya failed to answer.
Laya saw the whole room, not just what a camera near the player would show. These tests had no enemy bullets, bosses or items. In the small isolated tests, crossing any door counted as a navigation success, and staying alive until time ran out counted as survival. Neither meant clearing a floor. In the first Laya run, 3,590 of its 6,817 calls had only one possible answer, wait, so the raw call count overstates how many real decisions it made.
The wording and fight tests used saved moments from five of the development maps. Laya answered each moment three times per version, to check that it answered the same way each time. A wording change had to get more moments right, break none that were right before, and cause no errors. Every fight moment had at least one safe move. A fight change had to fix at least one damaging moment and break none. The wall facts also had to give the same answer every time and reproduce the original results exactly. No version passed. Changing option names or order also reached 28 of 32 exploring moments, but reversing the order fixed three and broke two, so I didn’t count those as wording fixes.
For the skill test, the bot had to win at least eight of ten floors with no crashes. Some full-floor tests ran with changes that had failed the saved-moment tests. Running them didn’t make those changes safe. The dice test used one fixed sequence per map, could pick wait, and still took Laya’s top answer in fights. Its lower revisit and call counts are partly because runs ended early.
For Laya to count as worth using, it had to beat random by at least 20 percentage points in wins, with 95% confidence. Or, with equal wins, it had to be at least 10% faster than random on four of the five random runs, again with 95% confidence that it was faster at all. Laya met neither bar. The statistics treated each map as one unit, kept random’s five runs on that map together, and compared times map by map rather than dividing averages.
Only a version that passed both bars would have run on ten fresh maps I had held back. It would need eight wins there too, with no tuning on those results. Those maps were never used, because Laya didn’t pass. The saved-moment tests also held back moments from the other five maps, for any change that passed. None did. Later full-floor tuning used those same maps, so those held-back moments are no longer fresh. The skill version also did not replace the short-action bot as the game’s default.
A later short-action test, under changed game rules, gave Laya 0/10, rules 9/10 and random 0/10. The short-action clip in this post was cut off, not run to the time limit. The current clips and the one current skill win don’t add to the September counts.
I started with a model choosing small movements and Rust helping it aim. I ended with the same model choosing goals and Rust working out how to accomplish them. The second bot won. Then random picked its goals, and it kept winning.