← All posts

Wasm SIMD Physics: How Web Games Handle 20,000 Colliders

Discover how WebAssembly SIMD supercharges browser collision detection, letting indie physics engines juggle 20,000 dynamic bodies at a silky 60 FPS.

Toss twenty thousand bouncing crates, ragdoll skeletons, or frantic arcade goblins into an HTML5 canvas, and older JavaScript engines will promptly turn your laptop fans into jet turbines. For years, browser-based physics tapped out after a few hundred intersecting shapes before frame rates tumbled off a cliff.

That ceiling has officially shattered. Thanks to mature browser support for WebAssembly (Wasm) SIMD (Single Instruction, Multiple Data), web-based physics engines are crunching millions of vector maths operations in the time it takes a standard JavaScript engine to clear its throat.

Whether you are building the next viral bullet-heaven swarm game or an utterly chaotic physics destruction sandbox, here is how SIMD-accelerated WebAssembly transforms browser collision detection from a laggy bottleneck into pure computational sorcery.


What Is WebAssembly SIMD in Web Physics?

Direct Definition: WebAssembly SIMD is a browser standard that exposes native 128-bit vector hardware registers to code compiled into Wasm (such as C++, Rust, or Zig). Rather than evaluating maths operations one variable at a time (scalar execution), SIMD enables engines to run identical mathematical operations across four 32-bit floats simultaneously in a single CPU instruction cycle, dramatically reducing the computational overhead of collision checks.

In browser game development, collision engines divide work into two phases:

1. Broadphase: Quickly eliminating objects that are nowhere near each other using Axis-Aligned Bounding Boxes (AABBs) or bounding spheres.

2. Narrowphase: Calculating exact contact points, manifolds, and restitution for bodies whose boundaries actually intersect.

Wasm SIMD shines brightest in the broadphase, where the engine evaluates tens of thousands of repetitive bounding-box comparisons every single frame.


Why Plain JavaScript Choked on Swarm Collisions

Classic JavaScript physics engines hit two brick walls when dealing with extreme object counts:

  • Memory Layout & Cache Misses: JavaScript objects are scattered across memory heaps. Traversing an array of complex entities forces the CPU to constantly hunt through RAM for property pointers, causing devastating cache misses.
  • Garbage Collection (GC) Hitches: Constantly allocating and dereferencing collision pair tuples triggers the browser's GC mechanism, introducing micro-stutters right when screen action peaks.

Even when web developers packed coordinates into flat Float32Array buffers, scalar code still evaluated bounding planes individually:


[Check Box A Min X] -> [Check Box A Max X] -> [Check Box A Min Y] -> [Check Box A Max Y]

Repeat that across 20,000 bodies checking against a spatial hash, and your 16.6ms frame budget vanishes before the rendering pipeline even receives draw calls.


Vectorising the Broadphase: Four Checks for the Price of One

Wasm SIMD leverages 128-bit vector types (v128). Because a standard single-precision float occupies 32 bits, a single v128 register holds four floating-point numbers at once (f32x4).

Instead of asking whether Box 1 overlaps Box 2, Box 3, Box 4, and Box 5 one by one, SIMD loads the bounding dimensions of four separate entities into a single register and runs intersection maths in parallel.


Scalar Processing (Single Register):
[ X1 ] vs [ TargetX ] ===> 1 Op
[ X2 ] vs [ TargetX ] ===> 1 Op
[ X3 ] vs [ TargetX ] ===> 1 Op
[ X4 ] vs [ TargetX ] ===> 1 Op
Total: 4 CPU Cycles

Wasm 128-bit SIMD Vector:
[ X1 | X2 | X3 | X4 ] vs [ TargetX | TargetX | TargetX | TargetX ] ===> 1 Op
Total: 1 CPU Cycle

In compiled Rust or C++ running under Wasm, the engine can execute an AABB overlap test across multiple candidates with a single vector comparison:


// Conceptual SIMD comparison using core::arch::wasm32
use core::arch::wasm32::*;

#[inline(always)]
pub unsafe fn check_four_overlaps(
    box_min_x: v128, box_max_x: v128,
    target_min_x: v128, target_max_x: v128
) -> v128 {
    // Check if target overlaps on the X axis for 4 targets simultaneously
    let no_overlap_left = f32x4_lt(target_max_x, box_min_x);
    let no_overlap_right = f32x4_gt(target_min_x, box_max_x);
    
    // Combine masks: 0 means an overlap occurred
    v128_or(no_overlap_left, no_overlap_right)
}

By unrolling spatial hash searches into groups of four (or eight with dual-issue pipelines), broadphase rejection runs blisteringly fast.


Performance Comparison: JavaScript vs Wasm vs Wasm + SIMD

The gap between raw JavaScript and vectorised WebAssembly is the difference between a sluggish prototype and a console-grade browser experience.

Architecture / TechniquePeak Active Colliders (60 FPS)Memory OverheadCache LocalityGarbage Collection Pauses
Pure JavaScript (Object-based)~600 – 1,200Very High (Pointers)PoorFrequent
JS TypedArrays (Data-Oriented)~3,500 – 5,000ModerateModerateRare
Standard WebAssembly (Scalar)~8,000 – 11,000Low (Linear Memory)HighZero (Manual Heap)
WebAssembly + 128-bit SIMD20,000 – 25,000+MinimalOptimal (Packed)Zero

Real-World Engine Integration: How Developers Pull It Off

Indie web game developers and engine creators (such as the teams behind Rapier.js and custom Box2D-Wasm ports) rely on a cohesive architectural pipeline to keep the rendering loop smooth:

1. Flat Linear Memory: Game state lives entirely within Wasm's linear memory buffer. Colliders are packed sequentially as Structure of Arrays (SoA) to guarantee memory alignment for vector instructions.

2. Worker Offloading with SharedArrayBuffer: Physics ticks run inside a dedicated Web Worker. The physics loop writes transform matrices directly into a shared buffer, allowing WebGL or WebGPU renderers to pull instance data without main-thread serialization penalties.

3. Narrowphase Early-Outs: SIMD isn't limited to broadphase boxes. Engines use vector dot products to compute Separating Axis Theorem (SAT) tests on convex hulls, discarding non-colliding polygons in a fraction of the clock cycles previously required.


What This Means for Web Games

When you can simulate 20,000 dynamic bodies inside a single tab without dropping frames, browser game design changes entirely:

  • Massive Horde Mechanics: Bullet-heaven and survival games can throw thousands of genuine, physics-driven enemies at the player instead of cheating with simple distance checks.
  • Destructible Environments: Physics sandboxes can fracture massive brick walls into thousands of individual debris chunks that tumble, bounce, and settle believably.
  • Complex Ragdoll Arenas: Multiplayer browser brawlers can afford fully articulated skeletal colliders for dozens of players at once.

WebAssembly SIMD closes the performance gap between native desktop binaries and the browser sandbox. If your physics loops are still grinding along on scalar JavaScript calculations, it is time to pack your structs, fire up a SIMD-enabled Wasm toolchain, and let the vectors do the heavy lifting.

Thanks for reading. Browse more from the Wobblox blog, or jump straight into all 100 free games.