- Add Scene.set(module) so JerryScript can switch scenes via
require()'d {init, update, dispose} modules; wire it into
engineUpdate() and rework overworldscene.js/init.js to the new shape.
- Fix scriptManagerExecFile() never pushing its own file's directory
onto the require() dir stack, breaking relative require() calls from
a top-level entry script (e.g. init.js -> ./overworldscene.js).
- Cache sprite geometry across frames for uislider/uitab/uiframe
(dialogs/textbox/settings) and give uitextbox a per-page glyph cache
instead of one spriteBatchBuffer call per visible character. uiconsole
drops its alloc/free vertex buffer for a fixed-size one. uifps skips
rebuilding its label when the text hasn't changed.
- Add PSP_OPTIMIZATION_PLAN.md and a handful of UI/spritebatch unit
tests covering the new caching logic.
- Add Dolphin/PPSSPP emulator smoke-test scripts (+ Docker variants and
CI jobs) for GameCube/Wii/PSP.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
15 KiB
PSP Optimization Plan
A survey of concrete memory and CPU (especially floating-point) optimization
opportunities for the PSP target, done 2026-07-31 on branch ac2 at commit
e61914ba + this session's scripting work. Point-in-time findings, not a
substitute for reading the referenced source before acting on it.
Why this exists
The PSP (Allegrex MIPS CPU, 222-333MHz, single core, 32-64MB main RAM, 2MB
eDRAM, 16KB I-cache / 16KB D-cache, no virtual memory or memory protection)
is the tightest-constrained target this engine ships to. Two rules guide
everything below, both already correctly applied in one place in this
codebase (entityposition_t caching its transform matrices instead of
recomputing them from position/rotation/scale every read — see "Already
correct" below):
- Memory is finite and cannot be compacted. No VM means a fragmented heap after a few minutes of play can fail an allocation even with "enough" total free bytes. Prefer static/pool allocation over malloc/free churn; prefer trading memory for CPU only when the memory cost is bounded and paid once.
- CPU is finite and floating-point math is not free, especially trig/sqrt on the plain scalar FPU. Prefer caching a computed result behind a dirty flag over recomputing it unconditionally; prefer avoiding a sqrt via squared-distance comparison wherever only a yes/no or ordering result is needed.
Priority 1 — Entity/component memory: a union tax paid on every slot, times 4 scenes
The single biggest finding. entitymanager_t (src/dusk/entity/entitymanager.h:11-15)
is a flat, statically-sized struct:
typedef struct entitymanager_t {
entity_t entities[ENTITY_COUNT_MAX]; // 64
component_t components[ENTITY_COUNT_MAX * ENTITY_COMPONENT_COUNT_MAX]; // 64*16 = 1024
componentid_t entitiesWithComponent[COMPONENT_TYPE_COUNT * ENTITY_COUNT_MAX];
} entitymanager_t;
component_t (src/dusk/entity/component.h:69-72) holds a tagged union
(componentdata_t) sized to its largest variant. That variant is
entityrenderable_t's spritebatch payload (entityrenderable.h:25-66):
spritebatchsprite_t sprites[64] at 40 bytes each (vec3 min + vec3 max + vec2 uvMin + vec2 uvMax) = 2560 bytes, versus ~28-300 bytes for every
other component type (camera, physics, animation, position, trigger).
Because the union is embedded by value in a flat array — not behind a
pointer, not sparse — every one of the 1024 component slots in every
entity manager costs ~2580 bytes, whether it holds a 28-byte camera or
nothing at all. 1024 * 2580 ≈ 2.52MB per entitymanager_t, confirmed
by the engine's own startup log every run: Entity manager size: 2684992 bytes (2622.06 KB).
scene_t (src/dusk/scene/scene.h:14-18) embeds entitymanager_t by
value too, and SCENE_COUNT_MAX = 4 (scene/scenebase.h:11) with
SCENE_MANAGER declared as a plain global (scene/scene.c:20) — so this
~2.52MB exists as static BSS from process start for all 4 scene slots,
used or not. Total: ~10.1MB reserved permanently for entity storage
alone — 16-31% of PSP main RAM — before a single texture, mesh, or audio
asset is loaded.
Options to fix (roughly increasing effort/risk):
- Pull
entityrenderablespritebatch_tout of the component union entirely — store spritebatch entities via a pointer/handle into a separate, smaller fixed pool sized to how many spritebatch-renderable entities actually coexist (almost certainly far fewer than 64 per entity, and far fewer entities need spritebatch at all vs. mesh/material rendering). This alone would shrink the union's dominant variant from ~2560 bytes to whatever the next-largest variant is (~300 bytes,entitytrigger_t) — an ~88% reduction, dropping the ~10.1MB down to roughly ~1.2MB. - Reduce
SCENE_COUNT_MAXfrom 4 if 4 concurrent scenes were never a deliberate requirement (check with the project owner — this may just be a round-number default nobody revisited), or make scene storage pointer-based/lazily allocated so unused scene slots cost ~0 instead of a fullentitymanager_t. - Reduce
ENTITY_COMPONENT_COUNT_MAX(16) if entities realistically use far fewer distinct component types simultaneously — check actual usage acrossentityprefablist.h/gameprefablist.hprefabs.
Do (1) first — it's the highest-leverage, most contained change (touches
entityrenderable.h's data layout and its render/dispose paths, not the
general entity/component system), and re-measure via the same startup
log line before deciding whether (2)/(3) are still worth doing.
Priority 2 — VFPU is completely unused; all vec/mat math is scalar
The PSP's Allegrex CPU has a VFPU (vector floating-point unit) capable
of fast SIMD-style 4-wide float ops and hardware-accelerated matrix
operations, exposed by pspsdk's pspvfpu/GU "Geometry Utility" (gum_*)
helpers. This engine links pspvfpu (cmake/targets/psp.cmake) but
never calls into it — confirmed via a zero-hit grep for
vfpu|pspmath|gum_|vcst|gu_matrix across src/duskpsp/ and src/duskgl/.
All vector/matrix math instead goes through cglm (glm_vec3_*,
glm_mat4_*), which auto-detects SIMD only for x86 (CGLM_SSE2/AVX) or
ARM (CGLM_NEON) — neither applies to MIPS, and no CGLM_* macros are
set anywhere in this repo (confirmed zero hits). Every glm_mat4_mul,
every position/rotation rebuild, every physics vector op runs on the
plain scalar MIPS FPU with the VFPU sitting idle.
This is the single largest available CPU win identified in this survey
— larger than any specific hot-path fix below, because it's a multiplier
on all of them. Concretely: route entityposition.c's matrix
rebuild/multiply path and physicsworld.c's per-body vector math through
pspvfpu/gum_* on the PSP build specifically (behind #ifdef DUSK_PSP,
matching the project's existing platform-guard convention), while keeping
cglm as the portable fallback for Linux/Knulli/GameCube/Wii. This is a
genuinely large effort (new platform-specific math backend, careful
correctness verification since VFPU has its own quirks around
pipelining/hazards) — scope it as its own project, not a quick pass.
Priority 3 — UI/spritebatch rebuilds and redraws every vertex, every frame
Already flagged in STATUS.md/ROADMAP.md (item 4/5) as a known open
issue; this session's research adds concrete numbers. meshvertex_t
(display/mesh/meshvertex.h:14-17) is { float uv[2]; float pos[3]; } =
20 bytes/vertex, no packing. SPRITEBATCH_SPRITES_MAX = 512,
SPRITEBATCH_FLUSH_COUNT = 16 → 32 sprites/flush = 192 vertices = 3,840
bytes rebuilt and redrawn per flush, with a flush forced on every
shader/material change (spritebatch.c:38-50) and used unconditionally
by every widget except uiconsole.c (uitab.c:56, uislider.c:183/199/231,
final flush in ui.c:51). A UI screen with ~50-100 sprites and 2-3
material changes costs an estimated 8-15KB of vertex rebuild + GU
submission every single frame, regardless of whether the UI changed
since the last frame.
Confirmed on the PSP legacy-GL path specifically: there's no GPU buffer
re-upload cost (meshFlushGL is a literal no-op comment: "we use the
glClientState stuff" — meshgl.c:91-93; meshDrawGL just points
glVertexPointer at the CPU-side array each call), so the real cost is
(a) the CPU-side per-sprite vertex rewrite every frame and (b) GU
re-transforming the full vertex stream from RAM on every draw with no
skip for unchanged geometry.
Fix direction (already captured in ROADMAP.md's debt backlog item 5):
extend uiconsole.c's cached-mesh, dirty-flag-gated rebuild pattern to
the rest of the widget tree. This plan adds one refinement: also consider
packing meshvertex_t down (next item) since it multiplies this cost.
Priority 4 — All-float vertex format wastes bandwidth and T&L cost
meshvertex_t uses two float fields (20 bytes) for every mesh and
sprite vertex, mesh-wide, with no normal or per-vertex color (color is
already handled at the material level, so that part is already optimal).
GU natively supports fixed-point 16-bit positions/UVs
(GU_VERTEX_16BIT/GU_TEXTURE_16BIT), which would roughly halve
per-vertex size and the corresponding vertex-fetch/transform cost, at the
cost of position precision (fine for UI/sprite work and most world
geometry at this engine's scale; worth checking against the largest
world coordinates actually used before committing). Pair with Priority 3
so the packing benefit compounds with the dirty-tracking benefit instead
of just reducing the cost of a rebuild that still happens every frame.
Priority 5 — Running gameplay logic in JerryScript has a real per-frame cost on PSP
This is new since last session's work, not a pre-existing issue: assets/ scripts/overworldscene.js's update() is now called every frame via
scriptManagerCallGlobal("update") (wired into engine.c's
engineUpdate()), replacing what used to be a direct cosf/sinf C
callback (entityUpdateAdd). Per call, scriptManagerCallGlobal
(scriptmanager.c:91-144) does: a global-object property lookup (key is
cached, but the jerry_object_get dispatch still runs), full JS
interpreter call dispatch (frame setup, argument marshaling) for what's a
two-line function, and an unconditional jerry_value_is_promise check
every call even though update() is synchronous. Inside, Math.cos/
Math.sin run as JS builtin calls rather than direct libm calls.
Additionally: this project's JerryScript fork is already patched to use
32-bit float internally (JERRY_NUMBER_TYPE_FLOAT64=0,
Findjerryscript.cmake) — a deliberate, already-correct optimization for
this exact concern — but the public engine API (jerry_value_as_number())
still returns double, so every native↔JS boundary crossing (every
moduleBaseArgFloat, every moduleBaseVec3ToObject) still pays a
float→double→float round trip. Also note JERRY_MATH is off, so
Math.sin/cos fall through to whatever generic libm the platform
provides rather than JerryScript's own fdlibm implementation — worth
checking whether turning it on changes anything measurable on PSP.
This is one update() call for one scene-level script today — cheap in
absolute terms. It becomes a real problem only if the pattern scales:
giving many individual entities their own per-frame JS update callback
would multiply all of the above per entity. Recommendation: keep
high-frequency, hot per-entity logic (movement, camera math, physics
response) in native C update callbacks (entityUpdateAdd, the existing
mechanism), and reserve JS for one-time setup, infrequent/event-driven
logic, and coarse-grained per-scene orchestration — which is exactly
what require() and the typed component wrappers built this session are
suited for, not a per-entity-per-frame hot path. This is a design
guideline to apply going forward, not a regression to fix in the current
overworldscene.js (one scene-level update() call per frame is fine).
Priority 6 — No pooling/arena allocator; plain malloc/free everywhere
memoryAllocate/memoryFree (util/memory.c) are direct malloc/free
passthroughs (memoryAlign→memalign, memoryReallocate/memoryResize→
realloc), with the only extra behavior being a global allocation-count
tracker used solely for test leak detection. No pool, arena, free-list,
or size-class allocator exists anywhere.
This survey found no rogue per-frame heap allocations in the hot
paths checked (render dispatch in entityrenderable.c, all of
physics/*.c) — so this is not an active bug today. But it's a
structural risk for a no-VM platform over a long play session: any future
code that does frequent small alloc/free (dynamic lists, string
building, temp buffers) will fragment the 32-64MB heap with no OS-level
recovery mechanism. Recommendation: before adding any new subsystem
that allocates/frees frequently at runtime (not just at load time),
default to a fixed-size pool or arena for it, following the same
"static, bounded" philosophy already used for ENTITY_COUNT_MAX/
SCENE_COUNT_MAX/ASSET_ENTRY_COUNT_MAX elsewhere in the engine, rather
than reaching for memoryAllocate per-instance.
Already correct — don't touch without new evidence
entityposition_t's matrix caching (the pattern the project owner called out as the reason for writing this plan). Verified: a full "dirty" recompute chain (decompose + rebuild local + rebuild world) is ~10 trig calls (asinf/cosf/atan2f) plus at most one 4x4 matrix multiply, andentityPositionEnsurePRS/EnsureLocal/EnsureWorld(entityposition.c:596-669) all early-return on a flag check, only doing that work when something actually changed. The cache is correct and clearly worth its ~128-byte-per-entity matrix storage cost.- Physics narrow-phase sqrt avoidance. All three
sqrtfcall sites inphysicstest.c/physicsshapemesh.calready sit behind a squared-distance early-reject and only compute the real (linear) distance once, when actually needed for penetration depth — no further "use squared distance instead" opportunity was found here. - No allocations in the render/physics hot loop (Priority 6) — this is good and worth preserving as new code is added to those files.
- Asset decompression already runs off the main thread. Assets are
DEFLATE-compressed (not stored) in
dusk.dsk, so loading pays a real CPU cost for decompression — but every*LoaderAsyncfunction (asset/loader/*/*.c) runs viaASSET.loadThread(assertNotMainThreadinasset.c:365), so this cost is already kept off the main thread. Low priority to change; if load-time CPU cost becomes a measured problem later, revisit stored-vs-compressed as a build-time flag rather than assuming compression is free.
Suggested approach
- Measure before changing. None of the above have been profiled on real PSP hardware in this pass — this is a source-level survey, not a profile. Before investing in Priority 1 or 2 especially, confirm with an actual PSP build/run (or at minimum the existing "Entity manager size" startup log plus a frame-time counter) that these are the real bottlenecks, not just the largest numbers on paper.
- Priority 1 first — it's the most contained (one component's data layout), has the clearest before/after metric (the startup log line), and doesn't require new platform-specific code.
- Priority 3 next (extend the console's dirty-tracking pattern to
the rest of the UI) — already scoped in
ROADMAP.md, no new design needed, just implementation. - Priority 2 (VFPU) as a dedicated project, not a quick pass — it's the largest potential win but touches core math plumbing and needs careful correctness verification on real hardware.
- Treat Priority 5 as a standing design guideline for all future scripting work, not a one-time fix.