Files
Toju/agents-docs/LESSONS.md
T
myxelium 718f4a99f0
Queue Release Build / prepare (push) Successful in 28s
Deploy Web Apps / deploy (push) Failing after 10m14s
Queue Release Build / build-windows (push) Failing after 15m12s
Queue Release Build / build-linux (push) Successful in 43m41s
Queue Release Build / finalize (push) Skipped
Queue Release Build / build-android (push) Successful in 17m25s
fix: AppImage required --no-sandbox
Not tested enough but works on my machine
2026-08-14 03:51:23 +02:00

73 KiB
Raw Blame History

Agent Lessons

Durable rules for AI agents working on this project.

How to use this file

At session start: read agents-docs/LESSONS-INDEX.md only. Open lesson bodies here only for tags that match the task. Do not load this entire file into context by default.

During the session: if the user corrects you, reverts your edit, or re-prompts with the same instruction — record a lesson here and add a one-line entry to LESSONS-INDEX.md before closing the task. See triggers in agents-docs/AGENT_WORKFLOW.md.

Format of a lesson: every entry uses the four-slot template below. Brevity matters — if you can't state the rule in one sentence, the lesson isn't sharp enough yet.

### <short imperative title>

- **Trigger:** what you were about to do that turned out wrong (one line, concrete enough to pattern-match against)
- **Rule:** what to do instead (one sentence, imperative voice)
- **Why:** the consequence of getting it wrong — past incident, hidden constraint, user preference
- **Example:** one concrete instance, ideally a code or command snippet

Keep lessons sharp. Tag each rule with one or two tags in square brackets after the title (e.g. [testing] [migrations]) so future agents can grep for relevance. If a rule no longer applies, delete it — stale rules drown the real ones.


Lessons

app.commandLine.appendSwitch cannot disable the Chromium sandbox [electron] [packaging] [linux]

  • Trigger: a packaged Linux build shows only the window background and spams Unable to access(W_OK|X_OK) /tmp / Creating shared memory in /tmp/... failed, while the same build works when the user types --no-sandbox.
  • Rule: sandbox and Ozone switches only count when they are on the real command line at process start. Never pair a runtime appendSwitch('no-sandbox') with appendSwitch('disable-dev-shm-usage') — the first is a no-op because the zygote has already forked, the second takes effect and redirects shared memory into /tmp, which the still-active sandbox denies forever.
  • Why: electron-builder's AppImage AppRun is a bash script that execs $APPDIR/<executableName> "$@" and ignores the bundled desktop entry, so linux.executableArgs never reaches a double-click or terminal launch. Only an installed .desktop file passes those arguments.
  • Example: electron/app/linux-launcher.rules.ts generates the launcher that tools/after-pack.js installs in place of the real binary (renamed <name>-bin); it enables --no-sandbox only where unprivileged user namespaces are denied (Ubuntu 24.04+ AppArmor, hardened kernels), since an AppImage payload is mounted nosuid and cannot fall back to the SUID helper.

A hold-on-unknown rule needs every attach site behind it [voice] [webrtc] [realtime]

  • Trigger: replacing a strict media gate with "hold an established path when nothing confirms the peer", while some fast path still attaches the track without asking the rule.
  • Rule: route every attach site — including the pre-offer shortcut in createPeerConnection and every sibling media kind — through the same decision, and refresh all of them on each piece of new evidence.
  • Why: an ungated attach becomes the guess the rule then protects: the track alone reads as an established path, so hold keeps sending indefinitely to a peer that never joined the channel. The strict gate used to erase that mistake on the next pass.
  • Example: MediaManager.mayOpenVoicePathToPeer() gates the first offer in create-peer-connection.ts, syncCameraRouting() reuses decideVoicePathRouting, and notePeerVoiceReport() calls refreshVoiceRouting() so a departure report reaches the camera too, not just the mic.

Missing gossip about a peer is not evidence it left voice [voice] [webrtc] [realtime]

  • Trigger: gating an outgoing media track on the observer's store copy of the remote user's voiceState, and detaching whenever that copy is absent.
  • Rule: close a negotiated media path only on positive evidence — we left voice, the peer itself reported leaving or another channel, or the connection is gone; treat absence as unknown and hold the path. Opening a path still requires confirmation, so a guess never starts sending the microphone.
  • Why: the signal server broadcasts user_left for any socket it declares dead, so a suspend or a flaky hop wipes that copy while the peer is still in the channel. The observer then detached its mic permanently — a silent member with no way back through the UI, not even by toggling mute.
  • Example: decideVoicePathRouting() in toju-app/src/app/domains/voice-session/domain/logic/voice-path-routing.rules.ts, used by MediaManager.syncVoiceRouting() for the mic and mayHearPeerVoice() for playback gain.

Assert continuity when the state you broke repairs itself [testing] [voice] [verification]

  • Trigger: proving a media cut by wiping a peer from the roster and then checking that audio still flows.
  • Rule: when the broken state is refreshed by a periodic message, sample the victim second by second across the window instead of asserting an end state.
  • Why: peers gossip their voice state every 5s (VOICE_HEARTBEAT_INTERVAL_MS), so the roster heals moments after the wipe and the mic re-attaches; the end-state check passed against the unfixed code and proved nothing. Only a per-second sample showed the peer losing audio.
  • Example: assertUninterruptedInboundAudio(peer, 10) in e2e/tests/voice/roster-loss-preserves-voice.spec.ts fails on the first silent second; the earlier assertTwoWayAudio after the wipe did not.

Never test suspend/resume against a live-reloading dev server [testing] [dev-shell] [verification]

  • Trigger: suspending the machine to check whether a voice call survives sleep/wake, with the windows served by ng serve.
  • Rule: disable the dev server's reload before any suspend/resume test (LIVE_RELOAD=false npm run dev), and treat a renderer reload in the results as an invalid run rather than a product finding.
  • Why: ng serve --ssl runs Vite over HTTP/2; the suspend destroys that stream, so on resume Vite throws The stream has been destroyed from viteTransformMiddleware into the error overlay of every window, and its live-reload client reloads the page. The reload re-bootstraps the app out of the call, so the post-resume readings showed a "connected" peer with zero RTP — which looks exactly like a silently dead call but only meant the reloaded app was no longer in voice.
  • Example: dev.sh appends --live-reload=false when LIVE_RELOAD=false; the first P7.4 attempt produced 20 audio stalled lines that proved nothing.

Keep diagnostic history outside the page you are diagnosing [testing] [verification]

  • Trigger: collecting samples into a window.__probe array in the DevTools console, then reading them back after the disruptive event.
  • Rule: persist probe samples to localStorage (or outside the renderer entirely) and stamp each sample with a per-load id, so a reload keeps the history and becomes visible evidence instead of silent data loss.
  • Why: the event under test is often the very thing that destroys in-heap state; a reload wiped every pre-suspend sample while leaving the old console lines on screen, so the probe looked loaded but __voiceProbe was undefined and the baseline was gone.
  • Example: tools/voice-probe.js stores samples under metoyou_voice_probe_v1 and reports RENDERER RELOADED when performance.timeOrigin changes between samples.

Record whether the user is in a call before calling zero RTP a failure [testing] [voice] [verification]

  • Trigger: asserting on inbound/outbound audio packets without also recording voice membership and local mic track state.
  • Rule: capture isVoiceConnected() and the local audio tracks' readyState in the same sample as the RTP counters, and only call a stall a stall when the client is supposed to be in voice.
  • Why: peer connections exist for chat data channels regardless of voice, so "connected with zero audio" is the normal reading outside a call; without the voice flag the two cases are indistinguishable and a healthy app looks broken.
  • Example: readLocalMedia() in tools/voice-probe.js logs in-voice mic=live, and the stall check is gated on current.voice === 'in-voice'.

Never answer second-instance by relaunching the app [electron] [dev-shell]

  • Trigger: making a second dev launch reuse the open window by restarting the running instance (app.relaunch(); app.exit(0)).
  • Rule: handle a second instance in place — focus and webContents.reloadIgnoringCache() — and never relaunch the process from the second-instance handler.
  • Why: the relaunched successor inherits the same dev argument and asks for the single-instance lock while the dying parent still holds it, so it is refused as yet another second instance and the pair respawns forever; every generation also exits 0 instead of the launcher's handoff code, so concurrently --kill-others tears down ng serve and the API server, and an in-flight loadURL dies as ERR_FAILED (-2) that reads like an unreachable dev server.
  • Example: resolveSecondInstanceAction() in electron/app/second-instance.rules.ts returns 'reload-existing', and deep-links.ts reloads instead of relaunching.

Never gate a presence indicator on the observer's own participation [ui] [voice] [webrtc]

  • Trigger: writing if (!isUserInCurrentVoiceRoom(...)) return false before reading a remote user's share/camera state.
  • Rule: decide a remote indicator from the observed user's state alone; keep the observer's own session out of the input entirely.
  • Why: a user sharing alone in a voice channel looked idle to everyone outside it, so nobody could tell there was anything to watch — while the peer plane had already delivered the announcement, because screen-state goes to every open data channel and not just voice participants.
  • Example: shouldShowStreamIndicator() in domains/voice-session/domain/logic/stream-indicator.rules.ts; guarded by e2e/tests/screen-share/outside-voice-live-indicator.spec.ts, where the observer never joins voice.

ERR_FAILED (-2) on a dev loadURL usually means aborted, not unreachable [electron] [dev-shell]

  • Trigger: blaming the cert or ng serve when Electron logs ERR_FAILED (-2) loading 'https://127.0.0.1:4200'.
  • Rule: read the rejection stack — stopLoadingListener means the navigation was stopped (window destroyed, app exiting), so look for whatever killed the process; SSL=true already appends ignore-certificate-errors.
  • Why: the cert and the dev server were fine; the app was exiting underneath the load, and chasing TLS wasted the first pass at the bug.
  • Example: loadDevelopmentClientWithRetry() in electron/window/dev-client-load.rules.ts retries and never throws, so the window still gets its listeners and shows a readable failure page.

Pin a chosen media device with deviceId: { exact }, never a bare string [webrtc] [media] [electron]

  • Trigger: the user picks a different microphone or camera and nothing changes — not mid-call, not after leaving and rejoining voice.
  • Rule: build getUserMedia constraints as deviceId: { exact: id }, and handle OverconstrainedError / NotFoundError by retrying once with the system default.
  • Why: a bare deviceId: id is an ideal constraint, so Chromium may satisfy it with the device it already had; the feature then looks broken while every unit test passes. exact makes the request fail loudly instead, which is why it needs the explicit fallback so an unplugged device degrades rather than killing the call.
  • Example: buildMicrophoneConstraints in audio-device-selection.rules.ts plus the single retry with SYSTEM_DEFAULT_AUDIO_DEVICE_ID in media.manager.ts captureMicrophone and direct-call.service.ts captureCallMicrophone.

A second dev Electron window needs its own --user-data-dir [electron] [dev-shell]

  • Trigger: launching a second desktop instance for a two-user test; the existing window blinks and reloads and no second window appears.
  • Rule: launch the peer with its own --user-data-dir (npm run dev:peer), and never launch the desktop shell from an agent shell.
  • Why: Electron's single-instance lock is scoped to the userData directory, so a default-directory launch hands its argv to the running instance instead; tools/launch-electron.js always appends --metoyou-dev-reload-existing, and the second-instance handler in electron/app/deep-links.ts answers that with app.relaunch(); app.exit(0). Separate data dirs are also what give the two windows separate identities.
  • Example: dev-peer.sh--user-data-dir="$DIR/.dev-userdata/$PEER_NAME".

An outage test that only re-checks the end state is not a guard [testing] [verification] [webrtc]

  • Trigger: writing or trusting a test that breaks something (kills a server, closes a channel), then asserts the feature works again afterwards.
  • Rule: also assert what must not have happened in between — for a call, that the RTCPeerConnection was never rebuilt (countCreatedPeerConnections unchanged) — and prove the assertion by temporarily injecting the regression.
  • Why: re-checking only the end state passes for a client that tore the call down and rebuilt it, which the user hears as a dropped call. Injecting peerManager.closeAllPeers() on signaling reconnect kept every audio and peer-count assertion green; only the connection-count assertion failed.
  • Example: e2e/tests/voice/recovery-preserves-media.spec.ts — "The call was never rebuilt behind the user back" compares counts captured before testServer.kill().

coturn hands out a relay candidate but refuses loopback peers by default [testing] [webrtc] [turn]

  • Trigger: a relay-only test (iceTransportPolicy: 'relay') where candidates gather fine but every peer connection ends up closed.
  • Rule: run a local coturn with --allow-loopback-peers (plus --log-file=stdout --verbose, or docker logs stays empty and readiness cannot be observed).
  • Why: without it coturn still allocates and Chrome still reports a typ relay candidate, so the failure looks like broken app code rather than a blocked relay; connectivity checks to the other 127.x browser are simply dropped.
  • Example: e2e/helpers/turn-server.ts--allow-loopback-peers next to --relay-ip=127.0.0.1.

Swap a live device with replaceTrack; an empty device list is missing evidence [voice] [webrtc] [devices]

  • Trigger: a settings picker changes a capture device (mic, camera) while a session is live, or code reacts to devicechange by re-reading enumerateDevices().
  • Rule: re-capture, then replaceTrack on the existing senders and stop the old track — never tear the session down and rejoin. Treat an empty (or id-less) device list as no information: only fall back to the system default when a populated list proves the saved id is gone. Ask for deviceId as a preference, not exact.
  • Why: voice-controls.component.ts called disconnect() then connect() for a mic change, so every peer saw a leave/rejoin and the user lost the channel; the settings pickers wrote localStorage and applied nothing. enumerateDevices() returns [] before microphone permission is granted and Firefox never lists audio outputs, so "not in the list" would silently reset a valid choice on startup. A plain track swap on an already negotiated sender needs no SDP exchange, so the swap is invisible to peers.
  • Example: MediaManager.switchInputDevice() + resolveAudioDeviceSelection() / buildMicrophoneConstraints() in domains/voice-session/domain/logic/audio-device-selection.rules.ts, owned by VoiceAudioDeviceService; proven by e2e/tests/voice/live-input-device-change.spec.ts (outbound audio keeps flowing, no rejoin broadcast).

One owner for a toggle the UI mirrors [voice] [state] [ui]

  • Trigger: two surfaces (in-channel controls and a settings modal, a tray and a window) each keep a local signal for the same boolean — mute, deafen, camera on.
  • Rule: keep the state where the effect happens and let every surface read it back through a computed; never reset a mirror to a hardcoded value on teardown.
  • Why: MediaManager owned isMicMuted / isSelfDeafened, but voice-controls.component.ts kept its own copies and reset them to false in disconnect(), so after leaving voice the button said unmuted while the track was still disabled — and playback was un-deafened behind the user's back.
  • Example: isMuted = computed(() => this.webrtcService.isMuted()) in voice-controls.component.ts; disconnect() passes the real state into voicePlayback.updateDeafened().

A timed-out sync round is not a clean one [messages] [realtime] [verification]

  • Trigger: deciding a poll/backoff cadence (sync, presence, reconciliation) from a timeout firing with nothing received, or from a fire-and-forget send that "asked" every peer.
  • Rule: model the round — who was actually reached, who replied, what they reported — and let only a fully answered round with nothing outstanding buy the slow cadence; re-arm the timer from the verdict of the round that just closed, never from the previous one.
  • Why: messages-sync.effects.ts set lastSyncClean = true inside syncTimeout$, so a round nobody answered dropped the poll from 10s to 15min; sendToPeer also returned void and only logged when the channel was closed, so peers listed in getConnectedPeers() (filled at connectionState === 'connected', before the data channel opens) counted as asked. On top of that, repeat({ delay }) read the flag at emission time, so a round that discovered missing ids was already committed to a 15-minute wait.
  • Example: message-sync-round.rules.ts (createInventoryRound / recordInventoryReply / isInventoryRoundClean) plus messages-sync.effects.spec.ts, which advances fake timers and asserts the fast cadence survives silence, a partial answer, an undelivered request, and a late reply reporting missing ids.

Derive a conversation id from canonical humans, never from the ids on the wire [direct-message] [identity]

  • Trigger: building or trusting a composite id (DM thread, call id, dedupe key) made of participant ids that arrived in a payload or came from a roster entry.
  • Rule: resolve every id through an alias index first (buildDirectParticipantAliasIndexgetCanonicalDirectConversationId / canonicalizeDirectConversationId), and collapse already-stored alias copies on first touch instead of only fixing new ones.
  • Why: getDirectConversationId sorted the raw pair, so a peer who addressed the local user by a provisioned foreign actor id produced a second thread; the recipient saw two conversations for one human and clicking the peer opened the empty one. Matching aliases for admission was already in place, which made the fork look like a delivery bug instead of an id bug.
  • Example: e2e/tests/chat/cross-signal-dm-identity.spec.ts fails with element(s) not found for the peer's message the moment the self-alias group is dropped from DirectMessageService.participantAliasIndex().

Report whether a call event was delivered before showing a live call [direct-call] [verification]

  • Trigger: calling a fire-and-forget send (sendCallEvent, broadcast, notify) and then moving the UI into the success state.
  • Rule: return the transport result, ring before joining local media, and surface "reached nobody" through the same error signal the view already renders.
  • Why: startCall joined voice first and dropped the boolean from PeerDeliveryService.sendCallEvent, so a call to an unreachable peer showed the caller in a live-looking session that would never connect.
  • Example: DirectCallService.ringParticipants sets deliveryError (call.errors.ringUndelivered) and private-call.component.ts folds it into callErrorMessage; the e2e drives it with window.simulateOffline() on the caller.

Never spend a retry budget on attempts the transport cannot deliver [realtime] [recovery]

  • Trigger: writing or reviewing a bounded retry loop (peer reconnect, resync, delivery) that counts attempts before checking whether the channel it needs is even available.
  • Rule: check the dependency first and defer without counting; spend an attempt only when it can actually reach the far side, and when the budget really does run out publish a state the UI can show and re-arm the loop when the dependency returns.
  • Why: peer-recovery.ts incremented reconnectAttempts before isSignalingConnected(), so a ~60s signal outage burned all 12 attempts doing nothing, then cleared the timer and deleted the tracker entry with no user-visible state and no re-arm — the peer stayed dead until an unrelated roster event happened to heal it.
  • Example: schedulePeerReconnect now defers while signaling is down, emits peerRecoveryStatus$ { status: 'failed' } at exhaustion, and resumeStalledPeerRecovery() re-arms from handleSignalingConnectionStatus.

Repair a dead data channel on the live connection before rebuilding the peer [realtime] [webrtc] [recovery]

  • Trigger: handling a closed/failed RTCDataChannel by tracking the peer as disconnected and rebuilding the whole RTCPeerConnection.
  • Rule: while the connection is still connected, have the deterministically elected initiator create a replacement channel on that same connection (no renegotiation needed — the SCTP transport is already up) and let the other side adopt the incoming channel; rebuild only as the fallback when the replacement never opens.
  • Why: the control channel dying took voice, camera, and screen share down with it, and replaceDataChannel was already implemented and wired but never called — the spec asserted not.toHaveBeenCalled() and the README described the soft replacement as if it shipped.
  • Example: e2e/tests/voice/recovery-preserves-media.spec.ts asserts the created-RTCPeerConnection count stays at 1 per peer after closeOpenDataChannels; forcing the rebuild path makes it fail.

Compare peer ids only within one signal server's identity space [realtime] [identity] [webrtc]

  • Trigger: about to compare a remote peerId / roster oderId against a local id — deterministic initiator election, offer-collision politeness, reconnect election, self-filtering, or the oderId stamped into a voice/camera/screen payload.
  • Rule: resolve the local id for that peer's signal server (getLocalOderIdForSignalUrl where the signalUrl is in hand, getIdentifyCredentialsForPeer inside the peer manager) and elect roles only through peer-role.rules.ts; never reach for the home credential.
  • Why: one human has a different actor id per signal server, so a home-vs-foreign comparison is not antisymmetric — both peers offer (glare) or neither does until the 5s takeover, which is the "some users can't hear each other" report. It also makes your own foreign roster entry fail the self-check, so the client tries to peer with itself.
  • Example: realtime-session.service.ts wired getLocalOderId to getIdentifyCredentials() (always home) while shouldInitiatePeer compared it against foreign roster ids.

Reproduce initiator/glare bugs with a simultaneous reconnect, not staggered joins [testing] [realtime] [webrtc]

  • Trigger: writing an e2e for peer election, glare, or "cannot hear each other" and joining clients one after another.
  • Rule: get every client onto the roster, then reload/reconnect them with Promise.all so all pairs elect from the same snapshot, and assert real audio flow plus exactly one initiator per pair.
  • Why: staggered joins let one side's 1s fallback-offer timer serialize negotiation, so a wrong comparison still converges and the test passes on broken code — three sequential-join runs passed against the known-bad wiring before the simultaneous reconnect made it fail on audio.
  • Example: e2e/tests/voice/cross-signal-initiator-election.spec.ts — 4 users, 2 home signal servers, one shared voice channel, Promise.all(reload).

Interview before coding; dont guess the fix [workflow] [bugs] [tokens]

  • Trigger: about to edit product code for a bug/feature after reading the ask or Obsidian note, while acceptance, approach, or scope is still ambiguous or has real alternatives.
  • Rule: send a short interview (understanding, gaps, A/B/C + recommended default, proposed scope, proof of done), wait for the users choices, then implement only that — skip only if they said “just fix it” / “no interview.”
  • Why: unprompted guesses cause wrong fixes and expensive back-and-forth; one clarifying turn costs less than a wrong implementation thread.
  • Example: fix bug "Images and files in chat doesn't load" → read the note → ask whether the failure is channel-switch blank vs cold reload vs both before touching attachment services.

Default to toju-app/ + targeted electron/ + CI; do not crawl the monorepo [workflow] [tokens] [scope]

  • Trigger: about to browse all of electron/, or to grep/Read under server/, e2e/, website/, or docs-site/ on a normal product bug without the user naming those packages.
  • Rule: stay in toju-app/, .gitea/workflows/, and only the Electron files on the renderer→preload→handler path; if the fix looks like server//e2e, ask once instead of exploring those trees.
  • Why: monorepo-wide (and whole-electron/) exploration multiplies context on expensive problem-solving models without fixing the asked client bug.
  • Example: attachment disk restore → toju-app persistence service + electron/preload.ts + the one IPC/file helper involved — not every file under electron/migrations/ or electron/api/.

Write HANDOFF.md and ask the user for a new chat — agents cannot open chats [workflow] [tokens] [handoff]

  • Trigger: the thread is long, the user says "handoff"/"new chat", or a new major objective starts while more work remains.
  • Rule: overwrite (never append) agents-docs/HANDOFF.md with Status: active and short sections, ask the user to start a new chat with that file; when the handoff task is finished, clear the file to Status: none with empty sections.
  • Why: fat chat history dominates token burn; an appending handoff file becomes a second fat archive that every new chat reloads.
  • Example: user: "handoff" → replace HANDOFF → reply: "Start a new chat and attach @agents-docs/HANDOFF.md." Later when done → reset HANDOFF to empty Status: none.

Prove the asked behavior; unit-green is not done [verification] [testing] [workflow]

  • Trigger: about to report a task finished because colocated Vitest specs (or a narrow mocked unit) are green, while the users ask was a product behavior, UI flow, or bug they can still reproduce.
  • Rule: treat acceptance as “the asked functionality works” — prove it with a user-visible path, focused e2e, or an explicit manual check; keep unit tests as support, never as the sole done signal.
  • Why: agents optimized for TDD often stop at implementation-shaped tests that pass while the real feature/bug remains broken, which wastes follow-up turns and burns tokens on false completion.
  • Example: for “DM reply doesnt show for the caller,” a passing DirectMessageService mock test is insufficient until the cross-signal conversation identity path is exercised (e2e or a behavior-level regression that fails on the old fork-thread bug).

Keep NgOptimizedImage off runtime blob and data URLs [angular] [images]

  • Trigger: Angular template lint suggests replacing [src] with ngSrc for a user-uploaded image rendered from blob: or data:.
  • Rule: Keep a plain src binding, document/disable prefer-ngsrc, and use native loading/decoding plus the app's own lifecycle controls; Angular throws NG02952 for blob/data ngSrc.
  • Why: NgOptimizedImage targets network/CDN images and cannot resize, preload, or safely manage renderer-created attachment blobs.
  • Example: chat attachment thumbnails use [src]="attachment.objectUrl" loading="lazy" decoding="async", never [ngSrc].

Read the exact Obsidian bug note before diagnosing a named ticket [workflow] [bugs]

  • Trigger: The user says fix bug "…", names a Bug - … ticket, or the worktree already contains plausible changes / a similarly named resolved ticket.
  • Rule: Resolve and read only that note under Log/Bugs/ (and its attachment folder if needed); use Expected Result as acceptance; fix in default scope; do not list the whole inbox or treat BUG_TRACKER.md's snapshot table as live.
  • Why: attachment reload-host changes looked related to “Images and files in chat doesn't load” but came from a separate resolved ticket and did not cover the reported channel-switch state regression; inbox-wide reads also burn tokens for no gain.
  • Example: fix bug "Images and files in chat doesn't load" → read /home/ludde/Nextcloud/Obsidian Vault/Log/Bugs/Bug - Images and files in chat doesn't load.md only, then implement against its Steps/Expected.

Run npm run i18n:sync after editing any public/i18n/catalog/*.json file [i18n] [testing]

  • Trigger: Added new call.errors.* keys to toju-app/public/i18n/catalog/call.json and used them in code; the full test run failed in app-i18n-catalog.rules.spec.ts with "Missing i18n keys" even though the keys existed in the catalog file.
  • Rule: The runtime and the catalog spec read the merged toju-app/public/i18n/en.json, not the per-area catalog/*.json files — after any catalog edit, run npm run i18n:sync (root script, tools/sync-app-i18n-catalog.mjs) and commit the regenerated en.json alongside the catalog change.
  • Why: without the sync the new strings silently fall back to raw keys at runtime and the catalog spec fails, but only in the full suite — targeted spec runs of the feature under change pass, so the failure surfaces late.
  • Example: npm run i18n:sync && npm run test after adding call.errors.microphonePermissionDenied to catalog/call.json.

Match direct-call recipients against every local identity alias, exactly like DMs already do [direct-call] [identity]

  • Trigger: "User receiving direct call doesn't get notified" — a caller who met the callee through a room on the caller's signal server addressed the ring by the callee's provisioned actor id; handleIncomingCallEvent admitted only payload.participantIds.includes(oderId || id), so the ring was silently dropped, the caller sat "In Voice", and the callee saw nothing. DMs had the identical bug fixed earlier (baa350e), but the fix stopped at DirectMessageService and never reached DirectCallService.
  • Rule: every self check on a cross-user event (admission, sender-echo filter, remote-participant filtering, DM-header peer lookup) must span all local aliases — home id, entity id, peer id, plus each SignalServerCredentialStoreService.listValidCredentials() actor id — and incoming aliases must be normalized onto the canonical local id before session state is keyed (normalizeDirectCallPayloadSelfAliases).
  • Why: the failure only reproduces when caller and callee have different home signal servers, which no same-server e2e covers; and when one identity-alias bug is fixed in a domain, grep for the same === currentUserId pattern in sibling domains that share the transport — the direct-call domain reused PeerDeliveryService but kept the naive check for another month.
  • Example: direct-call-participant-identity.rules.ts#directCallPayloadIncludesAnyId / normalizeDirectCallPayloadSelfAliases; regression e2e e2e/tests/voice/dm-header-call-ring.spec.ts registers Bob on a secondary signal server, meets in a primary-signal room, and asserts the DM-header call rings Bob's incoming-call modal (fails on old code, passes after).

Resolve outbound direct-call recipient ids to the peer's connected signal identity [direct-call] [identity] [signaling]

  • Trigger: cross-signal direct calls still failed after the inbound alias fix — the caller joined voice and showed "In voice" while the callee never rang. PeerDeliveryService.resolveSignalingPeerId returned null when the stored peer id was a home id but presence/route was registered under the provisioned actor id, so sendRawMessage was never called; even when attempted, the server relays only when targetUserId exactly matches the callee's connected oderId.
  • Rule: outbound DM/call delivery must collect every recipient alias (peer-delivery-identity.rules.ts#collectRecipientDeliveryCandidateIds), pick the routable id with pickRoutableRecipientId, always attempt signaling send (broadcast fallback when no single route works), and surface call.errors.recipientUnreachable to the caller when delivery cannot succeed — never leave the caller in a silent "In voice" state.
  • Why: inbound and outbound identity bugs are independent; fixing admission on the callee does not help if the ring never leaves the caller or hits the wrong targetUserId on the wire.
  • Example: PeerDeliveryService.sendViaSignaling + DirectCallService.resolveRoutableRecipientId; e2e e2e/tests/voice/dm-header-call-ring.spec.ts (callee-home room, people-search call).

Decide attachment receive admission once at request time; never re-gate size in the chunk handler [attachments]

  • Trigger: "Sending files between users doesn't really work" — a browser user clicked Request on a 1050 MB generic file, the request gate (canReceiveAttachment) admitted it for in-memory receive, the sender streamed chunks, but handleFileChunk still had a leftover hard size > MAX_AUTO_SAVE_SIZE_BYTES rejection on the in-memory path, so every chunk was dropped, no ack was ever sent, the sender's waitForAck timed out, and the GUI never changed.
  • Rule: canReceiveAttachment (request time) is the single admission decision; the chunk handler may only route between disk-streaming and in-memory assembly — any stricter size check there silently drops chunks the request gate already admitted.
  • Why: the failure is invisible in logs-from-the-outside: the sender's per-chunk sends look like a working transfer ("packages with size 32kb") until the ack timeout, and the receiver sets requestError only into memory that a re-request immediately clears — the user just sees a dead Request button.
  • Example: removed the MAX_AUTO_SAVE_SIZE_BYTES guard in attachment-transfer.service.ts#handleFileChunk; regression e2e e2e/tests/chat/large-generic-file-transfer.spec.ts sends an 11 MB .bin between two browser clients and asserts Request → progress → Download (fails on the old code, passes after).

Re-queue attachment auto-downloads on every message/room binding event; never trust one transport's ordering [attachments] [realtime]

  • Trigger: cross-user attachment sync e2e (chat-message-features.spec.ts) flaked ~50%: file-announce (WebRTC data channel) beat chat-message (signaling websocket) to the receiver, so the announce-time auto-download resolved roomId=null, silently gave up, and nothing ever retried — the receiver showed "Waiting for image source..." forever. A related bug: the stalled-download reset keyed only on "receivedBytes>0 && no pending request", but the pending-request marker is deleted on the first chunk, so any auto-download pass during an active transfer cancelled it mid-stream and the retry deadlocked against the sender's active-transfer dedupe.
  • Rule: events that complete the messageId -> roomId binding (chat-message in messages-incoming.handlers.ts) must call queueAutoDownloadsForMessage again — never assume file-announce arrives after the message, they ride different transports; and stall detection must gate on chunk-progress staleness (lastUpdateMs older than ATTACHMENT_STALLED_DOWNLOAD_THRESHOLD_MS), never on the absence of a pending-request marker alone.
  • Why: both halves fail silently (no error, no requestError set), so the UI just sits at 0 bytes; the flake is timing-dependent and invisible in single-client tests — only the two-client e2e with --repeat-each exposed it deterministically enough to fix.
  • Example: handleChatMessage now calls attachments.queueAutoDownloadsForMessage(message.id) after rememberMessageRoom; shouldResetStalledAttachmentDownload(attachment, hasPendingRequest, nowMs) in attachment-autodownload.rules.ts. Verified with npx playwright test -g "syncs image and file attachments|syncs multi-chunk" --repeat-each=4 (8/8 after, ~50% before).

Scope per-user UI state by user id, not by the client database [persistence] [multi-user] [custom-emoji]

  • Trigger: custom emoji "saved library" membership was a single savedByUser flag on the shared emoji row plus a long-lived singleton (CustomEmojiService) that merged state across logins — so a second account on the same client (and the Electron shared SQLite DB) inherited the first user's picker.
  • Rule: when state is "per signed-in user" but the asset/row store is shared (Electron custom_emojis, or a renderer singleton that survives logout), key the membership by user id in its own store (localStorage metoyou_custom_emoji_saved:<userId>, mirroring the existing per-user usage ranking) and rebuild it in loadForUser; never rely on a global row flag or assume the singleton was reset on logout.
  • Why: the browser already isolates rows per-user database, so the leak only reproduces in-session (no reload) and on Electron's shared DB — both invisible if you only test reloads; a row-level flag also can't represent two local users saving the same asset.
  • Example: CustomEmojiService.resolveSavedIds(userId, emojis) reads/seeds a per-user id set; e2e e2e/tests/chat/custom-emoji-user-binding.spec.ts runs the whole user switch in ONE page load (client-side router nav only) so the singleton-retention leak is actually exercised, and the second user joins the first user's server instead of creating one (in-session "create a second server" leaves sourceId empty and the submit disabled).

Don't strand signed-out mobile users on a logged-out dashboard [auth] [mobile] [routing]

  • Trigger: App.ngOnInit special-cased mobile — signed-out visitors landing on / or /dashboard were kept on /dashboard (the "login form has no mobile chrome" rationale), so mobile users got a logged-out dashboard and never saw a login screen on startup.
  • Rule: decide startup routing for signed-out users with the platform-agnostic pure rule resolveUnauthenticatedStartupRedirect(currentUrl) (auth-navigation.rules.ts) — non-public routes → /login (with safe returnUrl), public routes (/login, /register, /invite/...) → stay; do not branch on isMobile() here.
  • Why: the mobile exception directly contradicted the product expectation ("greet signed-out users with the login screen"); the login form already links to register, so there is no dead-end to avoid.
  • Example: unit auth-navigation.rules.spec.ts (resolveUnauthenticatedStartupRedirect('/dashboard') === { path:'/login', queryParams:{} }); e2e e2e/tests/mobile/mobile-login-on-startup.spec.ts sets a 390×844 viewport before navigating (so ViewportService.isMobile is true at bootstrap) and asserts /dashboard and / both land on /login.

"Shared from your device" must gate on local bytes, not uploader user id [attachments] [multi-device]

  • Trigger: a second device of the same user showed "Shared from your device" and hid the download affordance for a file uploaded from another device — isUploader(attachment) returned uploaderPeerId === currentUserId, but uploaderPeerId is the user id (set to currentUser.id in publishAttachments), so it is true on every device of the uploader, including ones that only synced metadata.
  • Rule: key the sharing/ownership UI off whether this device holds the bytes, not who uploaded it — use isSharingFromThisDevice(attachment, currentUserId) (= isUploaderUser && deviceHasLocalCopy) from attachment-sharing.rules.ts; deviceHasLocalCopy = available + blob objectUrl, or a non-empty savedPath/filePath (synced metadata strips local paths, so it correctly reads as "no copy").
  • Why: same-user devices do not P2P with each other and sync only via account_sync (which strips filePath/savedPath), so the second device legitimately has no bytes; claiming ownership blocked the only path to view/download. For the regression to even be reachable in e2e, account_sync's chat-sync-batch had to start carrying the attachments map (it previously dropped attachment metadata entirely) via pushSavedRoomMessagesViaAccountSync(..., loadAttachmentMetas).
  • Example: unit attachment-sharing.rules.spec.ts (isSharingFromThisDevice({uploaderPeerId:'u1', available:false}, 'u1') === false); e2e e2e/tests/chat/multi-device-attachment-sharing.spec.ts uploads on device A then logs device B in afterward so the account_sync_peer_online full-state push delivers the attachment, then asserts device B shows a Request button and no "Shared from your device".

Generate Android brand icons from the source mark; guard against stock Capacitor placeholders [mobile] [android] [assets]

  • Trigger: the Android app shipped the default Ionic/Capacitor launcher icon (and a white adaptive background) because no brand icon was ever generated into toju-app/android/app/src/main/res/.
  • Rule: regenerate launcher + splash from images/icon-new-rounded.png with npm run cap:assets:android (tools/generate-android-app-icons.mjs, uses sharp), set the adaptive background to brand purple #4A217A (never #FFFFFF), and have the adaptive icon reference @mipmap/ic_launcher_foreground PNGs (delete the stock drawable-v24/ic_launcher_foreground.xml vector). cap:sync is not needed — these live in the native project, not webDir.
  • Why: a native launcher icon can't be asserted through a browser, so the regression proof is a hash guard: mobile-android-launcher-icon.rules.ts records the SHA-256 of every stock placeholder and the tests fail if any density still matches one. Pixel checks (purple ring + white-cat centre) confirm the brand mark actually rendered.
  • Example: findStockCapacitorResources(hashByFile) must return []; unit mobile-android-launcher-icon.rules.spec.ts + e2e e2e/tests/mobile/android-app-icon.spec.ts (deterministic fs/pixel checks, no emulator).

Bind chat attachments to a pre-allocated message id, never by matching content [attachments] [chat] [mobile]

  • Trigger: caption-less media (videos/images sent with no text) grouped onto the message bubble above and left an empty message below on Android — ChatMessagesComponent dispatched sendMessage without an id, then a setTimeout re-discovered the message by entry.content === content (always '' for attachment-only sends) and called publishAttachments on it.
  • Rule: pre-allocate the message id in the component (planChatMessageSend in chat-message-send.rules.ts), dispatch it via MessagesActions.sendMessage({ id, ... }) (effect uses id ?? uuidv4()), and bind attachments to that exact id with publishAttachments(id, files) — never re-find the message by content/timing.
  • Why: empty content is shared by every attachment-only message, so content matching picks the newest match and races the async create-effect; on Android the create latency exceeds the old 100 ms timer, so the file binds to a stale sibling. The race is invisible on fast desktop browsers, so the deterministic regression proof is the unit test that asserts the dispatched action id equals the attachment-binding id, not an e2e timing game (see the "don't bump E2E timeouts for sync flakes" lesson).
  • Example: planChatMessageSend(...).attachmentBinding.messageId === plan.action.id enforced in chat-message-send.rules.spec.ts; behavioral guard in e2e/tests/chat/attachment-only-message-grouping.spec.ts (proves the id flows component→effect→attachment by requiring each caption-less attachment to render in its own bubble).

Attachment file persistence must be platform-agnostic, not Electron-only [attachments] [persistence] [mobile]

  • Trigger: AttachmentStorageService talked only to window.electronAPI, so canWriteFiles() returned false on Android (Capacitor) and in the browser — no bytes were ever persisted there, and after restart/logout-login the uploader hit "Your original upload could not be found on this device" / "no peer with this file".
  • Rule: keep the path/bucket layout in AttachmentStorageService but delegate raw IO to a pluggable AttachmentFileStore selected by PlatformService — Electron disk, Capacitor Directory.Data (lazy-loaded, inline media via convertFileSrc), and a per-user IndexedDB vfs for the browser with a finite maxPersistableBytes cap; gate transfer persistence on canStreamToDisk() / canPersistSize() so the cap degrades gracefully.
  • Why: the browser e2e harness can't test native disk, but the browser IndexedDB store is real persistence, so a single-client send → page.reload() → reopen-room test proves the whole persist/restore orchestration with no peer connected.
  • Example: attachment-file-store.ts + {electron,browser,capacitor}-attachment-file-store.ts; e2e/tests/chat/local-attachment-persistence.spec.ts waits for both byte records (vfs) and attachments records with savedPath (summed across all metoyou/metoyou::<user> DBs, since an empty anonymous-scope DB exists) before reloading.

Never count duplicate chunks toward transfer progress, and never finalize on byte counters [attachments] [webrtc]

  • Trigger: P2P attachments arrived corrupt everywhere ("only the first bytes") because concurrent auto-download triggers double-requested a file, the sender streamed it twice, and the receiver counted duplicate chunk deliveries toward receivedBytes — inflating it past size, which both dropped the remaining chunks (post-Security guard) and passed the receivedBytes >= size finalize shortcut over a sparse buffer.
  • Rule: in chunked transfer receivers, ignore an already-buffered chunk index entirely (no progress update), use dense buffers, and finalize only when every chunk index is present — never use byte totals as an alternative completion signal; dedupe streams on the sender per (messageId, fileId, peerId).
  • Why: byte counters lie as soon as any duplicate, retry, or concurrent stream exists, and sparse-array every/some skip holes, so "looks complete" checks silently pass on partial data (same trap as the custom-emoji sparse-array lesson).
  • Example: handleFileChunk / finalizeTransferIfComplete in attachment-transfer.service.ts; multi-chunk e2e coverage via expectMessageImageContentSha256 in e2e/tests/chat/chat-message-features.spec.ts (single-chunk files cannot catch assembly bugs — test with >64 KiB payloads).

Don't bump E2E timeouts for sync flakes - gate on presence and read server logs [testing] [realtime]

  • Trigger: a multi-client chat-sync E2E flaked on "message not visible" and the first instinct was to raise toBeVisible timeouts or add waits; the user correctly rejected this ("it's not a timeout issue").
  • Rule: when a cross-user E2E assertion flakes, first gate the assertion on an observable precondition (peer visible in the members panel), then diff the signaling-server logs of a passing vs failing run (joined server, user_joined, user_left, Removing dead connection) before touching any timeout.
  • Why: the flake was a server race — identify + join_server arriving in one TCP segment were processed concurrently, the join was dropped as unauthenticated, and room membership silently vanished; no timeout can fix a message that is never broadcast. Fixed by serializing per-connection message handling in server/src/websocket/handler.ts.
  • Example: failing run showed one joined server for Ludde then user_left on sibling-client close; passing run showed two. expectServerPeerVisible(page, displayName) in e2e/helpers/multi-device-session.ts is the presence gate.

When renaming an Angular route, sweep every navigate/url-match/doc reference [routing]

  • Trigger: the find-servers route was renamed /search/servers in app.routes.ts, but servers-rail.component.ts still called router.navigate(['/search']) (leave-server) and matched startsWith('/search') for the user-bar visibility signal, throwing NG04002: 'search' on leave and never showing the user-bar on the discovery page.
  • Rule: after changing a path: in app.routes.ts, grep the whole repo for the old literal (/search) across *.ts/*.html (router calls, startsWith/url-match signals) and docs (docs-site, .agents/skills/playwright-e2e/SKILL.md route tables, domain READMEs) and update them all in the same change.
  • Why: router.navigate to a non-existent path raises NG04002 and aborts navigation, and stale startsWith matches silently break route-derived UI state — neither is caught by the build (string literals) and there was no servers-rail spec to catch it.
  • Example: fixed isOnServers/router.navigate(['/servers']) in servers-rail.component.{ts,html}; canonical post-leave/discovery route is /servers (FindServersComponent), matching DashboardComponent's router.navigate(['/servers']).

Server discovery must fan out across all endpoints and self-heal on 404 — never hardcode a host capability blocklist [server-directory]

  • Trigger: the dashboard "Popular Servers" and /servers discovery view were empty for fresh users until they typed a search. The first fix added a static DISCOVERY_UNSUPPORTED_HOSTS blocklist (signal.toju.app / signal-sweden.toju.app) that short-circuited discovery to []; the production hosts later shipped the /featured + /trending routes (verified curl → 200 with servers), so the stale blocklist kept blocking exactly the default endpoints a fresh account has while ungated search still surfaced them.
  • Rule: discovery (getFeaturedServers/getTrendingServers) must fan out across getSearchableEndpoints() with forkJoin + deduplicateById (mirroring all-endpoint search), and detect capability at runtime — on a 404 from /api/servers/{featured,trending}, fall back per-endpoint to the public GET /api/servers listing (fetchPublicServerListForDiscovery) instead of returning []. Do not maintain a hardcoded list of hosts that "don't support" a route; it goes stale silently and the build can't catch it.
  • Why: legacy servers resolve /featured as /servers/:id and answer 404, so a 404→fallback keeps the default view populated everywhere without a blocklist; the empty-query view renders discovery sections (not search results), so any divergence between discovery and search makes it look broken while search works.
  • Example: fetchDiscoveryFromEndpoint + fetchPublicServerListForDiscovery in server-directory-api.service.ts; e2e/tests/servers/server-discovery-default.spec.ts proves a fresh account sees Popular Servers without searching AND that route-intercepting /featured+/trending to 404 still populates it via the fallback.

Server registration needs ownerPublicKey: oderId || id, and must not be fire-and-forget [server-directory] [rooms]

  • Trigger: creating a server appeared to work (the creator landed in the room view) but the server didn't exist on the backend — invite-link creation and search both 404'd. createRoom$ sent ownerPublicKey: currentUser.oderId with no fallback; on restored sessions oderId can be falsy (identify still works because it falls back to id), so POST /api/servers returned 400 Missing required fields, and the .subscribe() swallowed the error while createRoomSuccess fired regardless.
  • Rule: resolve owner identity as oderId || id everywhere it's required (the server rejects an empty ownerPublicKey), and give registerServer().subscribe() an error handler so a failed registration is never silent.
  • Why: verified against the live server — authed POST with a truthy ownerPublicKey → 201; authed POST with an empty one → 400; the swallowed 400 is exactly what produces a "ghost" room the creator can enter but no one can find.
  • Example: buildServerRegistrationPayload(room, currentUser, normalizedPassword) in toju-app/src/app/store/rooms/server-registration.rules.ts, used by RoomsEffects.createRoom$.

Identify must fall back to the legacy session token, not only the new credential store [realtime] [authentication]

  • Trigger: the multi-signal-server auth refactor changed resolveCredentialForSignalUrl to read only SignalServerCredentialStoreService; sessions restored from disk (and logins where user.homeSignalServerUrl is unset) have an empty credential store, so identify was skipped on every signal server ("Skipping identify because no session token is available") and users appeared alone — no presence, no peers, sent messages visible only to themselves. E2E never caught it because every e2e flow does a fresh register/login that writes the credential store directly.
  • Rule: when resolving the identify credential for a signal URL, prefer the per-signal credential but fall back to the legacy AuthTokenStoreService token reconstructed with the current home user's id/displayName; never gate identify solely on the new credential store.
  • Why: persistSessionToken always writes the legacy metoyou.authTokens store on login, but the per-signal credential store is only populated on fresh login (with a loginResponse) or successful migration/provisioning — so on reload it can be empty while a valid session still exists.
  • Example: resolveSignalIdentity(credential, legacyTokenEntry, homeUser) in signal-server-credential-resolution.rules.ts, wired through SignalServerAuthService.resolveCredentialForSignalUrl (which now passes this.authTokenStore.getTokenEntry(httpUrl) and a homeUser carrying id). Test cross-user behavior via a session-restore path, not just fresh login.

Keep the per-signal-URL identify credential resolvable from the store [realtime] [authentication]

  • Trigger: after the multi-signal-server auth refactor, SignalingManager.getLastIdentify was switched to getIdentifyCredentialsForSignalUrl, which only read an in-memory cache populated after identify() ran; a freshly (re)connected socket then emitted join_server before any identify and users silently never appeared in the presence roster (almost all multi-user e2e tests timed out waiting for the peer's room-user-card).
  • Rule: getIdentifyCredentialsForSignalUrl must fall back to resolving the credential from the credential store so a new socket's onopen re-identifies before it re-joins; never restrict it to only the in-memory identify cache.
  • Why: the server drops join_server/view_server on any unauthenticated connection, so an identify-less join is lost with no error and recovery only happens on a later reconnect (often beyond the 20s test timeout).
  • Example: server log showed join_server authed=false ... display=User dropped, then User identified: Alice on a different connection but no Alice joined server; fixed in signaling-transport-handler.ts by resolving via dependencies.resolveCredential(signalUrl) when the cache is empty.

Store clientInstanceId in sessionStorage not localStorage [realtime] [multi-device]

  • Trigger: same user logged in on two tabs, browsers, or synced profiles sees alternating "Disconnected from signaling server" and no cross-device chat/voice sync.
  • Rule: persist metoyou.clientInstanceId in sessionStorage (one id per tab/window) and clear any legacy localStorage copy on first read.
  • Why: server identify evicts stale sockets with the same (oderId, connectionScope, clientInstanceId) tuple; a shared localStorage id makes each client kick the other in a reconnect loop.
  • Example: ClientInstanceService.getClientInstanceId() writes to sessionStorage; two tabs get different ids and stay connected simultaneously.

Revalidate IndexedDB scope without reinitializing on every read [persistence] [performance]

  • Trigger: DatabaseService.ensureReady() called initialize() before every delegated read/write to fix user-scope races.
  • Rule: cache the last validated metoyou_currentUserId and only re-run backend initialization when that scope changes or an in-flight initialize completes with a different scope.
  • Why: per-operation revalidation fans out across ban lookups, room loads, and message reads, causing channel/chat UI to stay blank until repeated server clicks eventually win the race.
  • Example: ensureReady() returns immediately when isReady() and validatedUserScope still match getStoredCurrentUserId().

Restore local user scope before protected writes [authentication] [persistence]

  • Trigger: a logged-in in-memory user can create rooms or messages after metoyou_currentUserId was cleared by a late session-expired path.
  • Rule: before protected local persistence or server-directory actions, restore metoyou_currentUserId from the current user and avoid treating a live current user as unauthenticated.
  • Why: otherwise rooms/messages fall into the anonymous IndexedDB scope, and route checks redirect to login even though NgRx still has the authenticated user.
  • Example: MessagesEffects.sendMessage$, RoomsEffects.createRoom$, and server-directory create/join components call setStoredCurrentUserId(currentUser.id) before writing or joining.

Persisted local user state still requires a session token [authentication] [signaling]

  • Trigger: Users appear logged in from local storage but cannot see peers online or send chat after session-token auth shipped.
  • Rule: before connecting signaling or loading rooms for a persisted user, require a non-expired token in metoyou.authTokens; redirect to /login on SESSION_EXPIRED, auth_required, or auth_error.
  • Why: WebSocket identify is skipped without a token, so join_server, RTC relay, and presence never establish even though the profile exists locally.
  • Example: hasValidPersistedSession() in auth-session.rules.ts from loadCurrentUser$.

Declare MODIFY_AUDIO_SETTINGS for Android WebRTC mic capture [mobile] [android]

  • Trigger: Android users accept the microphone prompt but voice calls and channels still fail to join.
  • Rule: include android.permission.MODIFY_AUDIO_SETTINGS in toju-app/android/app/src/main/AndroidManifest.xml and preflight Capacitor capture through MobileMediaService.ensureVoiceCapturePermissions() before getUserMedia.
  • Why: Capacitor's BridgeWebChromeClient.onPermissionRequest requests RECORD_AUDIO and MODIFY_AUDIO_SETTINGS together; if the latter is undeclared, the combined grant is treated as denied even after the user taps Allow.
  • Example: ANDROID_REQUIRED_MANIFEST_PERMISSIONS in mobile-android-manifest-permissions.rules.ts.

Do not override Tailwind with box-sizing inherit [mobile] [css]

  • Trigger: mobile pages still overflow horizontally until devtools disables *, *::before, *::after { box-sizing: inherit } in global styles.
  • Rule: in src/styles.scss keep box-sizing: border-box on the universal selector (matching Tailwind preflight); never replace it with inherit from html.
  • Why: inherit overrides preflight and some nested component hosts resolve to content-box, so w-full plus padding becomes wider than the parent — especially visible on the mobile dashboard beside the servers rail.
  • Example: src/styles.scss @layer base universal rule uses border-box, not inherit.

Use the app-shell servers rail for mobile discovery pages [mobile] [layout]

  • Trigger: patching min-w-0 / overflow-x-hidden on the dashboard (or find-people/find-servers) while the page still renders wider than the phone beside an embedded servers rail.
  • Rule: on mobile discovery routes (/dashboard, /people, /servers, …) show the global app.html servers rail and render the page full-width in appWorkspace; keep embedded swiper+rail stacks only for chat/DM/call routes (shouldShowMobileAppServersRail in mobile-shell-layout.rules.ts).
  • Why: nesting a second rail+Swiper stack inside router-outlet fights the shell flex width and content keeps sizing to intrinsic width, clipping cards and inputs on every viewport.
  • Example: hideAppServersRail() in app.html + dashboard pageContent only (no local <app-servers-rail>).

Defer attachment blob hydration on Electron startup [attachments] [electron]

  • Trigger: fixing inline attachment display by eagerly calling tryRestoreAttachmentFromLocal() for every persisted attachment during initFromDatabase().
  • Rule: load attachment metadata at startup, but hydrate blob URLs only for the watched room on demand; read disk files through chunked IPC (readFileChunk) and yield between chunks/attachments so large images never block the renderer.
  • Why: restoring every saved attachment as a single base64 round-trip plus synchronous atob() can freeze Electron for seconds even after the shell paints.
  • Example: runInitFromDatabase() stops at loadFromDatabase(); restoreLocalAttachmentsForRoom() hydrates lazily via restoreAttachmentBlobFromDiskPath().

Lazy-load Capacitor modules on Electron/desktop [mobile] [electron]

  • Trigger: adding mobile facades that statically import Capacitor adapters or @capacitor/* plugins into shared Angular services used by the desktop app.
  • Rule: keep web/electron shells on web adapters synchronously and load Capacitor adapters/plugins only through dynamic import() after runtime === 'capacitor' — never top-level import '@capacitor/...' in code reachable from app.ts / DirectCallService.
  • Why: bundlers evaluate static Capacitor imports during Electron startup, which can freeze the renderer before first paint even when runtime detection would have chosen the web adapter.
  • Example: resolveMobileAdapter() in mobile-capacitor-adapter.rules.ts plus async capacitor-plugin-loader.ts / loadMetoyouMobilePlugin().

Use the upgrade transaction during IndexedDB schema migrations [persistence] [browser]

  • Trigger: bumping BROWSER_DATABASE_VERSION and opening existing stores via database.transaction(...) inside onupgradeneeded.
  • Rule: during onupgradeneeded, reuse event.transaction.objectStore(name) for existing stores and only call database.createObjectStore for missing ones — never start a second transaction while the version-change transaction is active.
  • Why: nested transactions abort the upgrade, authenticateUser storage prep fails, and login/register navigates before setCurrentUser so DM routes throw "Cannot use direct messages without a current user."
  • Example: ensureObjectStoreDuringUpgrade(database, upgradeTransaction, 'messages') in browser-database-schema.ts.

Wait for authenticateUser storage prep before post-login navigation [authentication] [browser]

  • Trigger: dispatching UsersActions.authenticateUser from login/register and immediately calling router.navigate(...).
  • Rule: wait for setCurrentUser or loadCurrentUserFailure (e.g. waitForAuthenticationOutcome(actions$)) before navigating to returnUrl or /dashboard.
  • Why: authenticateUser$ prepares per-user IndexedDB asynchronously; early navigation renders DM/shell routes before the current user exists in the store.
  • Example: await firstValueFrom(waitForAuthenticationOutcome(this.actions$)) in register.component.ts and login.component.ts.

Use dense arrays for chunked transfer buffers [custom-emoji] [webrtc]

  • Trigger: chunked P2P asset assembly marks a transfer complete after the first chunk because array.some() skips sparse holes created by new Array(total).
  • Rule: initialize chunk buffers with Array.from({ length: total }, () => undefined) (or another dense initializer) before using some/every/filter to detect completion.
  • Why: a single assigned slot in a sparse array makes .some((chunk) => !chunk) return false, so multi-chunk custom emoji transfers are dropped and peers never receive uploaded images larger than one chunk.
  • Example: CustomEmojiService.receiveTransferStart stores chunks: Array.from({ length: total }, () => undefined) instead of new Array(total).

Route custom emoji right-click through the native context menu [custom-emoji] [ux]

  • Trigger: adding a second emoji-specific context menu beside NativeContextMenuComponent, or attaching handlers only to <img> nodes.
  • Rule: mark emoji hosts with data-custom-emoji / data-custom-emoji-library plus data-custom-emoji-id, let NativeContextMenuComponent own add/remove actions, and use a capture-phase preventDefault so Electron/browser image menus do not override them.
  • Why: the shell context menu already intercepts every image right-click; duplicate menus fight each other and button/div wrappers miss img-only handlers.
  • Example: reaction pills and picker buttons carry the data attributes; resolveCustomEmojiContextMenuTarget() opens Add to emoji library / Remove from emoji library from the global menu.

Separate known emoji assets from saved library [custom-emoji] [ux]

  • Trigger: syncing remote custom emoji directly into the picker/library when it is first seen in chat.
  • Rule: store remote emoji as known renderable assets, but only show them in the user's picker after an explicit save action such as right-clicking the rendered emoji.
  • Why: users need messages to render, but they should control which seen emoji become part of their local emoji library.
  • Example: CustomEmojiService.emojis filters to saved emoji, while findEmoji(id) can still resolve unsaved known assets for message rendering.

Chunk custom emoji assets over data channels [custom-emoji] [webrtc]

  • Trigger: sending uploaded custom emoji image data through a single custom-emoji-full peer event.
  • Rule: stream custom emoji assets as a metadata envelope plus bounded custom-emoji-chunk events; use buffered sends for back-pressure, but never rely on buffering to make oversized messages safe.
  • Why: a single base64 data URL can exceed browser SCTP message limits and fire RTCDataChannel.onerror, breaking the app-wide chat channel.
  • Example: send { type: 'custom-emoji-full', customEmojiTransfer, total }, then custom-emoji-chunk events with small data slices.

Re-clear visible notification channels after recompute [notifications] [startup]

  • Trigger: fixing startup unread badges by only changing read-marker writes or initial hydration.
  • Rule: also check later loadMessagesSuccess and syncMessages recomputes, and re-clear the focused visible channel after applying derived unread counts.
  • Why: the startup-selected server can load or sync messages after it was marked read, reintroducing a channel unread badge even though the user is viewing that channel.
  • Example: NotificationsService.refreshRoomUnreadFromMessages(...) should clear activeChannelId for currentRoom after recalculating counts from a startup message batch.

Disambiguate nested chat cards [chat] [ui]

  • Trigger: removing a visual treatment from chat history when a system message has both an outer row wrapper and an inner pill/card.
  • Rule: preserve the intended inner timeline pill unless the user explicitly targets it; render system messages outside the themed chatMessageBubble wrapper and keep data-message-id off direct child divs.
  • Why: PM call-started history should stay as a compact centered pill, while theme CSS such as app-chat-message-item > div[data-message-id] can turn the full-width row around it into the unnecessary card.
  • Example: In chat-message-item.component.html, keep data-testid="chat-system-message" with rounded-full border bg-secondary/45, put appThemeNode="chatMessageBubble" only on the non-system branch, and place [attr.data-message-id] on the nested pill instead of the system row wrapper.

Use terminal Vitest when the test tool hangs [testing]

  • Trigger: VS Code test execution stays at "Starting test run..." without producing Vitest output.
  • Rule: run the focused spec through the terminal with cd toju-app && npx vitest run <spec-path> and report the direct Vitest result.
  • Why: the test integration can hang before starting the runner, while the terminal Vitest command returns quickly and gives actionable failures.
  • Example: cd toju-app && npx vitest run src/app/domains/game-activity/application/game-activity.service.spec.ts.

Do not add fake chrome around screenshots [website] [design]

  • Trigger: wrapping a real product screenshot in decorative titlebar/window chrome or placing oversized marketing headings beside copy without checking overlap.
  • Rule: use the screenshot's existing frame when it already includes app chrome, and top-align large heading/copy columns with explicit readable widths.
  • Why: duplicated chrome makes CTA/product previews look broken, and bottom-aligned large headings can cover accompanying text on the marketing site.
  • Example: website/src/app/pages/home/home.component.html should render the screenshot directly; host-section should use top-aligned heading and .host-section-copy columns.

Prefer npm run lint:fix over hand-fixing lint/format [verification] [lint] [tokens]

  • Trigger: about to manually re-indent, reorder imports, or tweak Prettier/ESLint-fixable style after seeing lint failures.
  • Rule: from repo root run npm run lint:fix (format + sort:props + eslint . --fix); only hand-edit remaining non-fixable errors.
  • Why: manual style fixes burn turns and tokens and often miss what the project script already auto-corrects.
  • Example: after code changes → npm run lint:fix → if exit 0, do not also rewrite imports by hand.

Verify lint exits 0 before claiming done [verification]

  • Trigger: about to report a task as complete after running tests but skipping ESLint.
  • Rule: run npm run lint:fix from the repo root (or npm run lint after fixes) and confirm exit code 0 before any "done" claim.
  • Why: npm run test only runs the toju-app Vitest suite — it doesn't cover the server, Electron, or website packages. ESLint (flat config in eslint.config.js) is the universal check across every package; type-style violations slip through tests and break Gitea Workflows for the next agent.
  • Example: npm run lint:fix && echo OK — only claim done after seeing OK. For Electron type errors specifically, also confirm npm run build:electron succeeds (it invokes tsc -p tsconfig.electron.json).

Use blob URLs for inline attachment previews [attachments] [electron]

  • Trigger: receiving users see broken image icons or video players that never start, but "Download" saves a valid file.
  • Rule: never bind attachment.objectUrl to file:// URLs for chat <img>, <video>, or <audio> — always create a blob: URL from the bytes on disk or in memory; keep savedPath/filePath for IPC download/open only.
  • Why: Electron runs with webSecurity: true, so renderer pages cannot load arbitrary file:// app-data paths even when CSP allows file:; IPC download still works because it reads the path in the main process.
  • Example: ensureInlineDisplayObjectUrl() in AttachmentPersistenceService, and URL.createObjectURL(blob) in finalizeTransferIfComplete / handleDiskFileChunk instead of getFileUrl(savedPath).

Resolve Electron drag-and-drop file paths with webUtils [attachments] [electron]

  • Trigger: large videos play after drag-and-drop upload, but after restart the uploader sees a peer-download error even though they sent the file from disk.
  • Rule: when accepting dropped or pasted files in Electron, call webUtils.getPathForFile(file) from preload (getPathForFile on electronAPI) and annotate the File before publishAttachments; never rely on File.path in the renderer.
  • Why: Chromium removed direct File.path access in modern Electron; without getPathForFile, large uploads only exist as in-memory blobs and cannot be copied into app data for reload playback.
  • Example: annotateLocalFilePath(file, { getPathForFile: electronApi.getPathForFile }) in ChatMessageComposerComponent.addPendingFiles.

Preserve uploader local attachment paths across sync [attachments] [persistence]

  • Trigger: large Electron uploads play from filePath after send, but after reload the uploader sees "The connected peers do not have this file right now" and must P2P-download their own file.
  • Rule: never persist synced attachment metadata with filePath/savedPath stripped — merge with stored local paths, finish attachment DB init before applying sync batches, and try local disk restore before sending file-request to peers.
  • Why: P2P sync intentionally omits local-only paths; a startup race can overwrite the uploader's saved filePath with null, and large videos (>10 MB) are not auto-copied to app data so only the original path can restore playback.
  • Example: copy large Electron uploads into app-data on publishAttachments, mergeAttachmentLocalPaths(incomingMeta, storedRecord) in persistAttachmentMeta, await persistence.whenReady() in registerSyncedAttachments, and tryRestoreAttachmentFromLocal() before any file-request.