2026-07-12 — the exception is closed: Detox + a real Android runtime
The one gap the mobile suite could not honestly cover (below, 2026-07-11) is now covered where the claim
is actually made: on a device. targets/mobile/e2e/ + .detoxrc.js build the release APK from the
Expo-CNG native project and drive it with Detox on an emulator (local: the rap_phone AVD; CI: the
android-device job on an Ubuntu runner with KVM).
What it asserts, and why it can't be gamed. Not "1000 prompts in under N ms" — that benchmarks
whatever box CI got that morning. The invariant is scaling, measured against a same-session baseline:
roll 20, then roll 1000 (50× the rows) on the same device, and compare. A virtualized list holds a
window of rows, so 50× the data must not mean 50× the memory or a collapsed frame rate. Swap FlashList
for a .map() and both explode. The numbers come from the platform's own accounting — adb shell dumpsys gfxinfo (total/janky frames, p50/p90/p95/p99) and dumpsys meminfo (real PSS) — not from a
stopwatch inside the test, so the app cannot flatter itself. Emulator noise (software GPU, CPU
contention) is identical in both rolls and cancels out.
It also asserts the app produced all 1000 (the wait is on "1000 generated", so a re-introduced cap
times out rather than passing quietly) and is still interactive afterwards.
Why release, not debug: a debug build runs the dev bundle with dev-mode React checks on. Its numbers are pessimistic fiction, and the promise is about what a user runs.
Why nothing native is committed: android/ is generated by expo prebuild (CNG) from app.json +
the @config-plugins/detox plugin, and is gitignored. There is no hand-edited native file to rot.
Run it: npm run test:mobile:device:build (once, or after a native-config change) → npm run test:mobile:device. It is deliberately not in npm test — it needs a booted device and takes
minutes; the fast gate stays fast.
Where the wall-clock goes (so nobody has to wonder whether "CI is slow" is hiding a real defect): Gradle compiles the Android app from source ≈ 9 min, the emulator downloads/boots/installs ≈ 5 min, and the three rolls themselves ≈ 2.5 min (20 → 13 s, 200 → 30 s, 1000 → 87 s). A run that takes much longer than that is a failing run sitting in its timeouts — which is exactly what happened before the harness bugs were fixed, and it briefly got mistaken for the app being slow.
The one number that is genuinely open: the engine costs 23–30 ms/prompt on the emulator against
0.16 ms/prompt in Node. The render is fine (272 ms for 1000 rows, flat memory) — the phone's whole
cost is the engine. Hermes has no JIT and the CI CPU is emulated, but ~150× is more than that should buy.
Tracked in next-steps.md; measure on a real phone before theorising.
Where it actually runs, honestly stated. The release APK builds locally (verified: 27 MB app +
2.9 MB androidTest, BUILD SUCCESSFUL), but the local Windows AVD cannot run it: the moment the app
renders, the emulator's own graphics stack dies —
surfaceflinger aborts inside GoldfishMapper::readFromHost (mapper.ranchu.so), which takes
system_server with it, and every subsequent adb install fails with a StorageManagerService NPE or a
broken pipe. Reproduced under all three GPU modes (swiftshader_indirect, host,
angle_indirect); it is the emulator/host graphics stack, not the app (the same APK is what CI
installs). So the gate lives in CI's android-device job (Ubuntu + KVM + a google_apis image),
which is the Android runner the exception always called for. When a healthy local AVD is available the
same two commands run it unchanged.
Two Windows landmines worth not re-hitting:
MAX_PATH. RN's new-architecture C++ codegen writes object files whose paths embed the full source path; underC:\Users\…\projects\random-ai-prompt\targets\mobile\node_modules\…ninja dies with "Filename longer than 260 characters". Build from a short path (subst R: <repo>→ build underR:\targets\mobile). Linux/CI is unaffected.- The NDK. AGP fetches the exact pinned NDK (~700 MB) — hours on a home line. The app compiles no
C++ of its own beyond RN's codegen, so any same-major NDK works:
gradle/ndk-override.init.gradle(opt in withANDROID_NDK_VERSION), a no-op in CI.
2026-07-11 — the mobile test suite (a11y · visual · perf) + one honest exception
The mobile target now has the same rigor as the web, driven through its react-native-web export
(playwright.mobile.config.js + tests/e2e-mobile/, served by scripts/serve-mobile-web.mjs). Ten
projects: 5 device sizes × both colour schemes (the app defaults to "system", so a light-only run
never renders the dark canvas the design is built around).
- Accessibility (
accessibility.spec.js) — axe, 60 tests. Zero serious/critical WCAG 2 A/AA. It found the app was, in screen-reader terms, largely unusable: 119 of 121TouchableOpacityhad noaccessibilityRole, so react-native-web emitted plain<div>s carryingaria-label(invalid —aria-prohibited-attr), and every icon-only control had no accessible name (button-name, critical). jest could not have seen this: RN a11y props only become ARIA once react-native-web renders them. - Visual baselines (
visual.spec.js) — committed per surface × size × scheme, with the rotating suggestion and all animation pinned so a diff means a layout change. - Performance (
perf.spec.js) — typing cost and pane-mount cost, both of which the browser reproduces faithfully because they're pure JS work.
First, a correction: the documented numbers are NOT limits
1000 prompts / 100k gallery / 100k-line editor are the levels the app supports with no performance loss. They are a promise about behaviour, not a cap, and the app never limits the user — past them it degrades gracefully rather than refusing. If it handles 1000 smoothly it will handle far more before trouble starts.
Writing the max-load test surfaced that the code had forgotten this in two places, and I initially made it worse:
| Where | The bug |
|---|---|
targets/web/frontend/lib/home/buildRoll.js |
MAX_PROMPTS = 50 — every web roll was silently truncated to 50. The app could not produce the 1000 prompts it documents, and a user asking for 200 got 50 with no explanation. |
targets/mobile/screens/GenerateScreen.js |
Clamped to 1000 in five places (settings, roll, both steppers, and the new field) — a mobile-only limit neither the engine nor the web had. |
The web's prompt-count <input> |
max={50}, capping the spinner and marking anything higher as invalid. |
All removed. The floor (≥ 1, integer) stays — that's validity, not a limit. And the tests had
enshrined the caps (expect(len(999)).toBe(50) // capped), which is how it survived: a test that
asserts a bug is the bug's best defender. They now assert the opposite (len(5000) === 5000).
The exception: the 1000-prompt max-load test is SKIPPED, and here is the evidence
The advertised ceiling (1000 prompts/roll) cannot be honestly verified in the browser proxy. In the RN-web export, 1 prompt renders instantly, 100 doesn't finish in ~2 minutes, and 1000 times out. That looks like a serious app defect, so it was measured rather than guessed:
| What | Result |
|---|---|
| Engine, 1000 prompts (nodeLoader) | 158 ms, linear (0.2 ms/prompt) |
| metroLoader (the loader mobile actually uses) | identical — 0.2 ms/prompt |
| Mobile's generate path | ONE setResults for the whole batch — no per-prompt re-render |
So neither the engine, nor mobile's content loader, nor the screen's state handling is the cost. What's
left is the renderer: @shopify/flash-list's WEB implementation does not recycle the way the native
one does, so the export mounts the whole batch. That is a property of the proxy, not of the
Android app.
This matters more than the test does: "fixing" the app to make that assertion pass would be optimizing for a renderer the app never ships on. So the test is kept, skipped, with the measurements inline — not deleted (which would pretend the promise is verified) and not left failing (which would cry wolf in the gate).
The 1000-prompt promise is therefore still UNVERIFIED on mobile, and can only be verified where it is actually made: on a device/emulator. That needs a Detox + Android-emulator harness on an Android runner, which measures real frame timings and cannot lie about device behavior. Flagged to the owner as the one exception in the mobile testing mandate.
2026-07-04 — large-scale performance suite (tests/perf/)
A dedicated Playwright suite guards the officially supported maximum simultaneous load (100k-image
gallery + 1000 prompts / ~10k images + a 100k-line Manage file, all at once). It runs against the real
release server (targets/web/backend/serve.js) via playwright.perf.config.js — so the Manage file-read and
fs.watch hot-reload paths are exercised for real; the 100k gallery feed + image bytes are route-mocked
in-spec. Serial (workers: 1) so scenarios don't skew each other's timing; Chromium launched with
--enable-precise-memory-info for the heap-ceiling checks.
- Specs:
gallery.perf(100k images, bounded DOM + smooth scroll),generate.perf(1000 prompts roll out + smooth scroll + fast Single↔Generate switch),manage.perf(100k-line list: windowed rows, responsive 100k-entry filter, entry↔raw switch),hotreload.perf(add/modify a 100k-line file on disk → no freeze, auto-refresh), andtabs.perf(the combined max load: all three loaded, tab switching + round-trip scroll quality + heap ceiling). - Robust, not flaky (by design): the primary assertions are structural — rendered DOM-node counts
(the virtualization proof) and a JS-heap ceiling — plus generous response/frame budgets
(
tests/perf/helpers.js#BUDGETS) set to flag pathology (an un-virtualized surface janks + blows the heap by orders of magnitude), not micro-noise. Fixtures/helpers intests/perf/fixtures.js+helpers.js; on-disk fixtures areperf-harness-*(git-ignored, removed on teardown). Supporting unit tests:targets/web/tests/lib/windowRange,targets/web/tests/providers/sharedSettings, and extendeduseImageBatches(instant placeholders + concurrency cap). - Run it:
npm run test:perf:scenarios(intest:alland theperf-scenariosCI job). Profiler:npm run profile(scripts/profile-scenarios.mjs) → DevTools traces +Performancemetrics + frame stats inperf-profile/(git-ignored). Note: the oldernpm run test:perfis the separate bundle-size budget — unrelated.
2026-06-29 — comprehensive coverage expansion
A full pass took the suite from "every test type represented" to "every module covered,
thoroughly, with valid/invalid/edge inputs" (plan: notes/plans/testing-coverage-plan.md).
419 headless tests now pass (Node 226 + SPA 193, up from ~138 + ~60).
- Engine (Node): direct unit tests for the
random*helpers, the loader-injected stages (listStore/list/block),blockManifest,promptFilesAndSuggestions,engineedges,settings/aliasesguards, extralistManifestcases, and a real-datanodeLoaderintegration test. - SPA (jsdom):
targets/web/tests/lib/**(keywords, manageTree, output, rewrite, online, sessionKeys, providerMeta, wrapperStore, dplInserts, dialects, useProvider),targets/web/tests/providers/**(transport, server/rewrite adapters, browser generate wrappers, dispatch + Netlify handlers),targets/web/tests/components/**(NsfwToggle, PromptResult, DplStatus, DplInsertBar, ProviderPicker, ApiKeyField). Network is mocked with MSW (targets/web/tests/msw/, wired intotests/setup.js,onUnhandledRequest: "bypass"). - Coverage gates (CI-enforced): root
vitest.config.jsthresholds — statements 88 / branches 76 / functions 88 / lines 90 (engine ~93% lines);targets/web/vitest.config.js— modest global floor + asrc/lib/**floor (lines/statements 65, functions 60, branches 50). CI now runstest:coverage(Node) andtest:coverage(SPA) so the gates actually fire.browserLoader.js,dpl/dplLanguage.js, andproviders/index.jsare excluded (covered via the browser/e2e path or not meaningfully unit-coverable). - Performance:
npm run test:perfbuilds the SPA and runsscripts/check-bundle-size.mjs(gzipped-JS budget, currently 763 KB vs a 900 KB budget);npm run test:lhciruns Lighthouse CI (lighthouserc.json, informational in CI — the bundle budget is the hard gate). - Cross-browser:
npm run test:e2e:all(setsPLAYWRIGHT_ALL_BROWSERS) runs the e2e + a11y specs on Chromium + Firefox + WebKit + a Pixel-7 mobile viewport; visual-regression stays Chromium-only. One-time:npx playwright install firefox webkit. - New CI jobs:
cross-browser(FF/WebKit/mobile, visual skipped) andperf(bundle budget hard gate + Lighthouse informational), alongside the existing check/targets/web/e2e jobs.
The reality (as of 2026-06-22)
The project now has a full automated test suite built on Vitest (Node + jsdom)
and Playwright (browser). It covers the active engine and the SPA across every
standard test type. The pre-revival CLI + classic server were removed from the tree, so they're out
of scope; the two live pipeline stages they once shared (cleanup.js, prompt-salt.js) now live in
engine/core/stages/ and are tested with the rest of the engine.
Layout
tests/ # Node-side suite (Vitest, environment: node)
helpers/ # seededRandom.js (mulberry32 + withSeed), fakeLoader.js
unit/ # pure-module unit tests
integration/ # engine pipeline over a fake loader
contract/ # (provider/API contracts live in gui — see below)
snapshot/ # seeded, reproducible output snapshots
regression/ # bug-regression guards (one per fixed defect)
e2e/ # Playwright specs: home.spec, visual.spec, accessibility.spec
# + visual.spec.js-snapshots/ (committed visual baselines)
vitest.config.js # root (node) config
playwright.config.js # builds the SPA + serves dist via `vite preview`
targets/web/
tests/ # SPA suite (Vitest, environment: jsdom)
*.test.js # unit (share, settings, customStore) + contract (providers)
*.test.jsx # component/UI (Field, TokenPicker) via Testing Library
promptEngine.integration.test.js # real browser engine over the bundled data
setup.js # jest-dom matchers + localStorage reset + RTL cleanup
vitest.config.js # jsdom config (reuses vite.config: react plugin, lodash alias)
Test types covered
| Type | Where | Notes |
|---|---|---|
| Unit | tests/unit/**, targets/web/tests/*.test.js |
contentSafety, diffSettings, keywordRepeater, gatedLists, listManifest, DPL compiler, cleanup, prompt-salt; SPA share/settings/customStore |
| Component / UI | targets/web/tests/*.test.jsx |
Field controls, TokenPicker — React Testing Library + jsdom |
| Integration | tests/integration/**, targets/web/tests/promptEngine.integration.test.js |
full stage pipeline via a fake loader (Node) and via the real bundled-data browser loader (SPA) |
| E2E | tests/e2e/home.spec.js |
Playwright drives the built SPA: type → generate → results; block search |
| Visual regression | tests/e2e/visual.spec.js |
toHaveScreenshot of stable chrome (topbar, sidebar, full page with the random suggestion masked) |
| Accessibility | tests/e2e/accessibility.spec.js |
@axe-core/playwright, WCAG 2 A/AA, fails on serious/critical (color-contrast excluded — tracked) |
| Snapshot | tests/snapshot/** |
seeded (Math.random) DPL + pipeline output |
| Contract / API | targets/web/tests/providers.test.js |
SD WebUI txt2img request/response contract, fetch mocked |
| Smoke | scripts/smoke-test.mjs (npm run smoke) |
the original import-graph smoke, still the fast gate |
| Bug regression | tests/regression/bugRegressions.test.js |
one guard per fixed defect / documented landmine |
Running
npm test # lint + smoke + Node unit/integration/snapshot/regression + SPA suite
npm run test:unit # Node-side Vitest only
npm run test:web # SPA Vitest only (jsdom)
npm run test:e2e # Playwright E2E + visual + a11y (builds the SPA, serves dist)
npm run test:all # everything, including E2E
npm run test:coverage / test:web:coverage # with coverage
npm run test:e2e:update # refresh committed visual baselines
npm run test:e2e needs the Playwright browser once: npx playwright install chromium. The config sets
channel: "chromium" so the full chromium build is used (no separate chromium-headless-shell download).
Windows runtime prerequisite: Chrome-for-Testing needs the Microsoft Visual C++ Redistributable.
Without it, the browser fails to launch with spawn UNKNOWN → "side-by-side configuration is incorrect".
On this machine the bundled Chrome-for-Testing build hit the SxS error even with the VC++ runtime present,
so the config uses channel: "chrome" (the system-installed Google Chrome, version-matched to the
Chromium Playwright targets). CI can drop that channel to use the bundled browser. The Vitest suites
(npm test) have no browser dependency. First test:e2e run writes the visual baselines under
tests/e2e/visual.spec.js-snapshots/ (committed).
Gotchas baked into the suite
- lodash captures
Math.randomat import, so_.random/_.sample/_.shufflecannot be stubbed by overridingMath.random. Tests that touch lodash randomness assert invariants (token counts, value shape) or use single-entry lists; only the DPL renderer (its ownMath.random-based RNG) is made deterministic withwithSeed. - The SPA Vitest config reuses
vite.config.jssoimport.meta.glob(the browser loader's data bundle) and thelodashalias resolve exactly as in the real build. - Visual baselines are committed under
tests/e2e/visual.spec.js-snapshots/; regenerate them on a deliberate UI change withnpm run test:e2e:update.
Adding a bug-regression test
When you fix a bug, add an it("regression: …") to
tests/regression/bugRegressions.test.js that fails on the old behaviour and passes on the
fix, with a one-line note on the original symptom. That permanently locks the fix.