Status: Approved (2026-09-28), part of the v1 contract.
This document, together with manifest-v1.md, is the normative rendering contract for the
face lens manifest v1. The format is tagged format-v1.0.0 (engine 1) and format-v1.1.0
(engine 2, §11.8–11.12); a field-level change here needs the process in §10.9.
Screen effects have their own rendering contract, effect-rendering-v1.md.
Before the split this document also covered them, as lens kind 2. Those items moved there;
each keeps its number here as a one-line pointer, and no number is reused, so a citation of an
unchanged item still points at the same text.
format-v1.1.0 (engine 2). iOS SDK 0.2.0 shipped on format-v1.0.0, which is now fixed.
The animation upgrades — anchor and roll-follow tracks, custom curves, lanes, signal-driven clips
(§11.8–11.12) — are additive optional fields at engine level 2 (10.9), tagged format-v1.1.0.
Everything below this paragraph about "one version" is the history of format-v1.0.0.
One version until the SDKs ship (history). Before an SDK shipped, every change to v1 was
part of format-v1.0.0 itself: the tag was moved forward to include it. Every renderer (the editor preview's reference
renderer, the iOS and Android SDKs) follows the text as tagged. The edits made after the
first draft of this document are listed below so an SDK that pinned an earlier commit knows
what moved.
Clarifications — they pin what the first draft left open: 2.5 (no-face wording now
points at 6.10), 3.7 (a sprite hidden by endBehavior: 2, then shown or played), 4.3 (a LUT
grades un-premultiplied colour), 6.9 (order of events within one frame — clocks, particle
emission and expiry, the start event, host events, face triggers and emission start; emission
instants in whole microseconds), 7.7 (emission instants evaluated in whole microseconds, as
6.9 (1c) specifies) and 10.12 (the per-instance state list).
Behaviour decisions made after the first draft — not gap-fills: each changes what the
first draft literally required for some valid manifests. No SDK renderer existed when they
were made (both SDK repos were at scaffold plus the format pin), so no renderer had to change.
- (a) Triggers start armed (6.10; 2.5 and 6.3 point at it). Under the first draft's edge
wording (6.3) and 2.5's "face-signal values read as 0" without a face, a
threshold: 0
trigger never fired and a signal already high on the first frame was left undefined. Now a
threshold: 0 trigger fires on the frame a face is acquired or re-acquired, and a signal at
or above threshold on the first frame with a face fires.
- (b) Moved: the flash's exclusion from shake and pulse is a screen-effect rule
(
effect-rendering-v1.md E8.1, E8.3).
- (c) Face-particle lengths are fixed at emission (7.9; 1.6 amended). The first draft's 1.6 listed
particle
speed, gravity and sizeScale among the lengths re-evaluated every frame. Now a
particle takes W once, at emission, for every length.
Added before anything shipped: keyframe animation — the optional animations list and
action 6, play animation (§11, 7.11; 5.4, 6.9 steps 1a–1b, 10.6 and 10.12 amended). Like the
full-face anchor (10.11) it is engine level 1: part of format-v1.0.0.
manifest-v1.md defines the manifest's shape and validation. It does not say how a renderer
interprets the values. The iOS SDK, the Android SDK and the editor's web preview are three
independent renderers that must agree in golden tests, so every interpretation below has to be
pinned once, here.
Each item carries one marker:
- [S] Settled by spec (§x) — already decided in the design spec
(
2026-09-28-ingevora-mirage-design.md); cited, not re-opened.
- [P] Default — no field change — a rule on top of the existing fields.
- [D] Adopted with a new or changed field — see Decisions above (D1, D2); resolved.
#Decisions
Both decisions below were approved as recommended, adding required fields to v1 before it is
tagged (no formatVersion bump).
- D1 — sprite end behaviour. What a non-looping sprite (
sheet.loop: false, type 4)
shows after its last frame. sheet.endBehavior is now a required INTEGER on the shared
spriteSheet object: 1 hold last frame (the layer stays visible) · 2 hide (the layer's
visible becomes false; a later show or play sprite action makes it visible again from
frame 0, per 7.6 and 7.3). It applies when loop is false; when loop is true it is still
required (strict, simple) and has no effect.
- D2 — face-particle offset. Face particles (type 5) gain a required
offset: { x, y },
reusing the existing offset schema (x, y in [-2, 2], face widths, face frame — identical
meaning to transform.offset). Particles spawn at anchor point + offset, rotated with the
face frame exactly as a sticker offset is (2.6).
Everything else in this document is [S] or [P]. Note that the schema is strictObject
everywhere and engine levels exist only per layer type, anchor and event source, so any field added
after the tag — even an optional one — makes older SDKs reject the lens with no
minEngineVersion to explain why (see 10.9).
#1. Coordinate systems
#Face space
1.1 [S] §2, §4.2, §7.2 — A face lens is drawn into the outgoing frame; the processed frame
is what remote viewers receive, and the self-view shows those same processed frames.
1.2 [P] The renderer's input and output are the un-mirrored camera image, upright
(the host rotates the buffer, or passes its orientation, before rendering). Mirroring for a
self-view is a display transform the preview view applies after rendering; the sent frame is
never mirrored.
Rationale: one image space for all renderers; text on a sticker reads correctly to the people
who receive the video.
1.3 [P] Image axes: origin at the top-left pixel corner, +x right, +y down, pixel units
of the output frame. Pixel centres are at (i + 0.5, j + 0.5).
Rationale: matches every platform's 2D image convention (Core Graphics flipped, Android
Canvas, HTML canvas).
1.4 [P] "Left" and "right" in anchor names (2/3, 5/6, 8/9) are the subject's anatomical
left and right. In the un-mirrored image the subject's left eye appears on the image's right.
Each SDK maps its tracker's naming to this; the landmark fixtures (10.4) pin it.
Rationale: trackers disagree on naming; anatomical is the only unambiguous choice.
1.5 [P] Face frame. Let EL, ER be the subject's left and right eye centres.
M = (EL + ER) / 2 is the face origin. The face-frame x axis u = normalize(EL − ER) (points
to image-right for an upright face); the y axis v is u rotated +90° (points down, toward the
chin). Roll = atan2(u.y, u.x) in degrees: 0 for an upright face, positive when the head
tilts clockwise on screen.
Rationale: eye centres are the landmarks every tracker (Vision, ML Kit, MediaPipe) returns
most stably.
1.6 [P] Face width W = 2.2 × |EL − ER| in pixels (2.2 ≈ bizygomatic width ÷
inter-pupillary distance for an adult). All lengths — scale, offset, particle
speed, gravity, sizeScale — are multiples of W, re-evaluated every frame for
non-particle layers; particles take W once, at emission (7.9). (Amended after
the first draft: decision (c) in the header.)
Rationale: a face-contour width depends on each tracker's contour model and on yaw; the
eye-distance proxy is identical across trackers. Foreshortening under yaw shrinks W, which is
the intended 2D approximation of "follows head pose" (§5.3: v1 is 2D affine).
1.7 [S] §5.3 — v1 layers are 2D: anchors plus affine transforms. No yaw/pitch perspective,
no occlusion by the head.
1.8 [P] Sticker placement (type 2, and type 4 with placement.mode 1). The asset (or
one sprite frame) is drawn with its centre at P = A + offset, width scale × W, height
from the asset's aspect ratio (frame aspect for sprites).
followRoll: true — the offset is in the face frame (P = A + W·(offset.x·u + offset.y·v))
and the image is rotated by roll + rotationDeg about its centre.
followRoll: false — the offset is in image axes (P = A + W·(offset.x, offset.y)) and the
image is rotated by rotationDeg only.
Rationale: centre pivot is what an editor's handles imply; "follow roll" should carry the
offset with it, or a hat drifts off a tilted head.
1.9 [P] Rotation: degrees, positive = clockwise on screen (y-down), 0 = the asset
upright as authored. Applies to rotationDeg and directionDeg. directionDeg 0 = +x (right),
90 = down, −90 = up (the face-sparks fixture's −90 fires upward).
Rationale: consistent with the y-down image axes; matches CSS rotate() for the web preview.
1.10–1.12 — Moved: the overlay's space and its unit S are the screen effect contract's
(effect-rendering-v1.md E1.1–E1.3).
1.13 [P] Gravity is a constant acceleration along +y (down) in image axes, in
W/s²; negative values accelerate upward. It is never rotated by head roll.
Rationale: gravity is a world direction; a particle fired straight up with speed 0.8 and
gravity 1 rises 0.32 W and falls back.
#2. Anchors
2.1 [P] Measured anchors (from landmarks, per frame):
2 left eye = EL, 3 right eye = ER (eye centres).
4 nose = the nose-tip landmark.
7 chin = the lowest point of the chin contour (menton) — follows the jaw when the mouth
opens.
2.2 [P] Derived anchors (fixed offsets in the face frame from M, in W):
| Anchor |
Face-frame position (x, y) |
| 1 forehead |
(0, −0.30) |
| 5 left cheek |
(+0.22, +0.28) |
| 6 right cheek |
(−0.22, +0.28) |
| 8 left ear |
(+0.50, +0.10) |
| 9 right ear |
(−0.50, +0.10) |
| 10 full face (2.2a) |
(0, +0.15) |
(+x toward the subject's left, per 1.5.) Rationale: trackers have no forehead point and no
ears, and their cheek points differ; fixed offsets are identical everywhere. The constants are
proposals for the phase-0 spike to tune before the tag — after the tag they are contract.
2.2a [P] Full face (anchor 10) is the centre of the whole face, brows to chin,
at face-frame (0, +0.15) — halfway between a hairline at about −0.55 W and the menton at about
+0.86 W above/below M for an adult face. It is placed exactly like every other anchor (1.8):
with scale 1 the layer is one face width W wide, height from its image's aspect, rotated by
roll + rotationDeg when followRoll is on. Face particles on it spawn at that point (2.6).
It is a derived anchor, not the menton-based midpoint, so it does not move when the mouth opens.
The scenario packages/render/fixtures/scenarios/lens/full-face-mask.json pins it.
Rationale: a mask needs the face's centre and the face's size; both are already in the face
frame, so no tracker needs anything new. It is engine level 1 and part of format-v1.0.0
(added before any lens was published or any SDK shipped), so every engine-1 renderer (the SDKs'
1.0.0) implements it.
2.3 [P] Anchor positions are not smoothed by the renderer. Any landmark smoothing belongs
to the SDK's tracking stage, upstream of the renderer, and is off when face data is supplied
externally (§4.2 "external face data") — which is how golden tests feed faces (10.4).
Rationale: smoothing filters are the first place two SDKs would diverge.
2.4 [P] One face. If several faces are detected, the lens tracks the one with the
largest W; ties go to the one whose M is closest to the frame centre. The choice is
re-evaluated every frame.
2.5 [P] No face. While no face is tracked: face-anchored layers (types 2, 4-mode-1, 5
emission) are not drawn / do not emit, without changing their visible state; full-frame
layers (1, 3, 4-mode-2) keep drawing; already-emitted particles keep simulating; face-signal
triggers are not evaluated and re-arm (6.10; behaviour decision (a) in the header).
2.6 [D2] Face particles (type 5) spawn at anchor point A + offset, rotated with the
face frame exactly as a sticker offset is (1.8).
#3. Sprite sheets (type 4)
3.1 [P] Frames are laid out row-major from the top-left: columns = floor(sheetWidth / frameWidth), rows = floor(sheetHeight / frameHeight); frame k is the cell at column
k mod columns, row floor(k / columns). Leftover pixels on the right/bottom edges are
ignored.
3.2 [P] frameCount > columns × rows is invalid, rejected where images are decoded:
at upload/publish by the API, and by the SDK's validate-before-cache (§4.3 "every image
actually decodes"). Issue code asset.sheet_too_small (decode-time, not in checkRules,
because the manifest has no image dimensions).
3.3 [D1] Frame index during playback: k = floor(tPlay × fps / 1000) where tPlay is ms
since playback started (5.x). loop: true → k mod frameCount; loop: false → after
frameCount − 1, the layer follows sheet.endBehavior: 1 holds on the last frame · 2
hides (visible becomes false; a later show or play sprite shows it again from frame 0,
per 7.3 and 7.6).
3.4 [P] A sprite layer that is visible but not playing (never played, e.g. autoplay: false) shows frame 0.
3.5 [P] Sampling: bilinear, clamped to the frame's own cell (inset half a texel) so
neighbouring frames never bleed in; no mipmaps.
3.6 [P] Type 4 with placement.mode 2 and type 3 frame overlays have no fit: type 3 is
stretched to the full output frame (borders stay on every edge); type 4 mode 2 uses
cover (centred, aspect preserved, overflow cropped).
Rationale: a border must touch all four edges on every aspect ratio; an animation must not
distort.
3.7 [P] (clarification, D1 addendum) When a non-looping sprite with endBehavior: 2
passes its last frame — visible or not (a sprite hidden by an action keeps its clock, 7.2) — its
playback stops and visible becomes false. A later show makes it visible showing frame 0 without playing
(7.1 — show sets visible only; 3.4); a later play sprite replays it from frame 0 (7.3).
The end-hide is evaluated when the sprite clock advances at the start of a frame (6.9), so a
show or play sprite on the frame the sprite ends wins.
#4. Colour look (LUT, type 1)
4.1 [P] Layout as already documented (64³, 8 × 8 tiles). The LUT maps sRGB-encoded
input to sRGB-encoded output; no linearisation.
Rationale: the authoring tools that export LUTs work in encoded values; three renderers
agree trivially.
4.2 [P] Sampling is trilinear: r, g, b ∈ [0,1] map to lattice coordinates c × 63;
red and green interpolate bilinearly within a tile (texel centres), blue interpolates linearly
between the two neighbouring tiles.
4.3 [P] intensity blends in encoded space: out = mix(in, lut(in), intensity); alpha
is untouched; the LUT asset's alpha channel is ignored. A LUT grades un-premultiplied
colour: in is the pixel's straight colour, so its grade never depends on its alpha. A
premultiplied renderer divides by alpha before the lookup and multiplies the result by the
unchanged alpha; a pixel with alpha 0 stays as it is. (Clarification after the first draft.)
4.4 [P] A LUT grades everything below it in layer order (the camera image plus
lower layers) — not layers above it. Put it at index 0 to grade only the camera.
#5. Layer draw order and blending
5.1 [S] §3.2 — layers[] is ordered.
5.2 [P] Array order is back to front: index 0 is drawn first (nearest the camera
image / host UI), the last layer on top.
5.3 [P] Compositing is source-over with premultiplied alpha, in sRGB-encoded
space (no linear blending). PNGs are decoded as straight alpha and premultiplied once at load;
every asset is treated as sRGB and embedded colour profiles / gamma chunks are ignored.
Rationale: matches HTML canvas (the editor preview) and is cheap on the oldest devices
(§4.4).
5.4 [P] opacity (types 2, 3) multiplies the premultiplied source. Layers without
opacity (4, 5) draw at 1. An animated opacity (§11.4) multiplies the same way — for type 4 too,
resting at 1; face particles keep 1.
5.5 [P] Within a particle layer, particles draw in emission order (oldest first), each as
an upright quad centred on its position, width sizeScale × W, height by the asset's
aspect ratio, opacity 1 for its whole life, removed when its age reaches lifetimeMs. No
per-particle rotation or fade in v1.
5.6 — Moved: overlapping screen-effect instances (effect-rendering-v1.md E4.6).
#6. Triggers
6.1 [P] Face signals are scalars in [0, 1] computed per frame from the tracked face:
| Signal |
Value s (clamped to [0, 1]) |
mouth.open |
inner-lip gap (upper-lip-inner centre to lower-lip-inner centre) ÷ 0.35 W |
brows.raised |
(mean brow-centre-to-eye-centre distance ÷ W − 0.10) ÷ 0.08 |
blink |
1 − (mean over both eyes of lid gap ÷ eye width) ÷ 0.30 |
head.tilt |
|roll| ÷ 30° (either direction) |
smile |
(mouth-corner distance ÷ W − 0.36) ÷ 0.12 |
Rationale: each is a ratio of landmark distances, so it is scale-invariant and computable
from the same injected face data in every renderer. Constants are spike-tuned before the tag
and pinned by signal fixtures (10.4).
6.2 [P] Default threshold when omitted: 0.5 for every face signal.
6.3 [P] Face-signal triggers are edge-triggered: fire on the frame where s goes from
< threshold to ≥ threshold. The arm state in 6.10 makes this exact, including the first
frame and threshold: 0 (decision (a) in the header).
6.4 [P] Re-arm with hysteresis: after firing, the trigger fires again only after s
has dropped below max(0, threshold − 0.1). No time-based cooldown. A lost face (2.5) re-arms.
Rationale: stops one mouth-opening from firing on every jittery frame, without a hidden
timer.
6.5 [P] A trigger fires with s measured on the frame; its actions apply to that same
frame's render.
6.6 [P] Lifecycle: lens.started fires once, on the first frame processed after the lens
is attached.
6.7 [P] Host events are delivered to the attached face lens (running screen effects
receive them too, effect-rendering-v1.md E5.2); they are applied at the next rendered frame.
Events arriving before the lens has started or after it is detached are dropped, never
queued. Events with no matching trigger are ignored.
6.8 [P] When one event matches several triggers, triggers run in array order and actions
in array order within each, all before the frame renders; later actions win (show then
hide in the same frame = hidden).
6.9 [P] (clarification) Order within one frame. At frame time t:
- Clocks advance to
t, in this order:
a. Scheduled clip starts (§11.2): every autoplay clip whose startMs lens time has reached
starts, its clock at startMs. A face lens has no other scheduled starts (a screen effect's
are effect-rendering-v1.md E5.4).
b. Sprite clocks, including the D1 end behaviour (3.7), and the ends of non-looping clips
(§11.3).
c. Particle pools: continuous emission that was active at the end of the previous frame
(7.7) emits every due t_k ≤ t, in order — unless no face is tracked on this frame, in
which case emission stops and nothing is emitted for the interval (2.5). Before
each emission at t_k, the particles whose age at t_k has reached lifetimeMs are removed,
so the cap (7.4, 10.7) is evaluated at that instant. Then every particle whose age at t has
reached lifetimeMs is removed (5.5). Emission instants and particle ages are whole
microseconds, computed in IEEE-754 double precision and rounded to the nearest integer
with halves rounded up (every value here is ≥ 0; e.g. JS Math.round, Swift
.rounded(), Kotlin Math.round — not round-half-even): the frame's
T = round(t × 1000) (t in ms); T_0 is the T of the frame on which emission started
(step 5); T_k = T_0 + round((k × 1 000 000) / ratePerSecond) (multiply, then divide); a
particle's birth B is T_k for continuous emission and the frame's T for an
emit; t_k is due when T_k ≤ T; a particle has reached its lifetime at instant X
when X − B ≥ lifetimeMs × 1000; its age in 7.9 is (T − B) / 1 000 000 seconds.
Particles emitted for t_k < t spawn at this frame's anchor + offset, with this frame's
W.
- On the first frame only: autoplay (7.6), then the
lens.started triggers (6.6).
- Host events, in arrival order (6.7) — on the first frame this step is empty: host events
that arrived before the first frame was processed are dropped.
- Face-signal triggers, in array order.
- Emission start (7.7): every particle layer whose continuous-emission conditions now hold but
which is not emitting starts emitting with
t_0 = t, so its k = 0 particle is emitted now.
(A hide stops emission and clears the pool when it is applied, 7.2; so show then hide on
one frame emits nothing and draws nothing from the stream, 7.10.)
- Draw from the resulting state.
All actions of steps 2–4 apply to this frame's render (6.5), and later actions win across the
whole sequence; 6.8 still orders the triggers matching one event. Particles emitted by an action
(7.4) are born at t.
Example: on one frame a host event hides a sprite and a mouth.open trigger shows it. The host
event runs first (step 3) and the face trigger second (step 4), so the sprite is drawn.
Rationale: one fixed order is the only way two renderers agree when a host event and a face
trigger touch the same layer on the same frame; pinning emission and expiry to event instants
(t_k), not to frames, keeps particle pools identical at any frame rate.
6.10 [P] (behaviour decision (a) after the first draft, see the header; 2.5 / 6.3 / 6.4
addendum) Each face-signal trigger is armed
or disarmed, and starts armed. On a frame with a tracked face, an armed trigger whose signal
s ≥ threshold fires and disarms; a disarmed trigger re-arms when s < max(0, threshold − 0.1)
(6.4). On a frame with no tracked face, face-signal triggers are not evaluated and all re-arm.
This state machine is authoritative where it differs from the edge wording of 6.3 — notably a
trigger with threshold: 0 fires on the frame a face is acquired or re-acquired, and not while
the face is absent.
#7. Actions
7.1 [P] show / hide are idempotent: they set visible and nothing else. show on a
visible layer and hide on a hidden layer do nothing.
7.2 [P] hide on a particle layer clears its live particles and stops continuous
emission; hide on a sprite keeps its playback
clock running (visibility and playback are independent).
7.3 [P] play sprite implies show, and restarts from frame 0 if already playing.
Rationale: playing an invisible sprite is never the intent; restart is what "play" means on
every media API. (The face-crown fixture's show + play stays valid, just redundant.)
7.4 [P] emit particles implies show, then emits count particles at once. If fewer than
count slots are free under maxParticles, the excess is dropped — live particles are
never evicted.
Rationale: deterministic and allocation-free.
7.5 — Moved: run motion is a screen-effect action (effect-rendering-v1.md E6.5).
7.6 [P] autoplay: true (type 4) = an implicit play sprite at lens.started, without
changing visible. A hidden autoplaying sprite runs its clock and appears mid-animation when
shown.
7.7 [P] Continuous emission runs at ratePerSecond while the layer is visible and
autoEmit is true and a face is tracked. Emission
times are t_k = k / ratePerSecond for k = 0, 1, 2, … measured from when emission became
active (re-started at 0 on each show), evaluated in whole microseconds as 6.9 (1c) specifies. ratePerSecond: 0 never emits continuously.
7.8 — Moved: burstCount is a screen-particle field (effect-rendering-v1.md E6.7).
7.9 [P] Particle motion is evaluated in closed form, never integrated:
p(t) = p0 + v0·t + ½·g·t² with t the particle's age in seconds, v0 = speed × (cos θ, sin θ) and g = (0, gravity), in W units. W is taken at emission: particles live in
image space and do not follow the head afterwards. Every particle length — speed, gravity,
sizeScale — uses W taken at emission; 1.6's per-frame re-evaluation applies to non-particle
layers only.
(Behaviour decision (c) after the first draft, see the header.)
Rationale: identical positions on any frame rate.
7.10 [P] Spread and randomness. θ = directionDeg − spreadDeg/2 + spreadDeg × r,
r ∈ [0, 1). r comes from a mulberry32 stream per layer per instance, seeded with the
32-bit FNV-1a hash of the UTF-8 string "<lensId>:<layerId>", one draw per emitted particle in
emission order (dropped particles draw nothing). Angle is the only random quantity in v1.
Rationale: without a pinned PRNG, particle layers can never pass a golden test.
7.11 [P] play animation (action 6, §11.2) implies show of its clip's layer and
(re)starts the clip from 0, replacing any clip playing on that layer; it cancels that clip's pending
autoplay start. It names a clip (animation), not a layer.
#8. Timing
8.1 [S] §3.2 — A face lens has no durationMs and runs until stopped (a screen effect
declares one, effect-rendering-v1.md E7.3).
8.2 [P] Clock. Lens time = the input frame's presentation timestamp minus that of the
first processed frame (never wall clock). Golden tests supply timestamps directly.
8.3 — Moved: startMs is a screen-effect field (effect-rendering-v1.md E7.2).
8.4 — Moved: the durationMs cut and effect.ended (effect-rendering-v1.md E7.3).
8.5 — Moved: a host stopping an effect early (effect-rendering-v1.md E7.4).
8.6 [P] A face lens's lifecycle ends when the host detaches it; there is no end event.
Its triggers stop, and state is discarded (re-attaching starts fresh and fires lens.started
again).
8.7 [S] §4.4 — On sustained overload the engine drops the face lens (raw frames pass
through) and reports it.
8.8 — Moved: moving effect events between devices (effect-rendering-v1.md E7.5).
#9. Screen motion — moved
Screen motion is a screen-effect layer: effect-rendering-v1.md §E8 (shake, flash, pulse, and how
they combine). Items 9.1–9.4 are not reused.
#10. Other gaps a renderer needs pinned
10.1 [P] Output size: a face lens outputs exactly the input frame's size and pixel
format; it never crops or letterboxes.
10.2 [P] Asset colour/format: every asset is an 8-bit sRGB PNG (RGBA or RGB); 16-bit
or palette PNGs are expanded to 8-bit RGBA at decode. Texture dimensions ≤ 1024 are already
enforced at upload (manifest-v1.md notes).
10.3 [P] Sticker aspect: width is scale × W; height follows the asset's (or frame's)
own aspect ratio — never the face's.
10.4 [P] Golden-test inputs. Add (outside the manifest) fixtures/render/: each case
= a manifest + a sample frame + an explicit face-data file (the landmark points this document
uses: eye centres, nose tip, menton, inner-lip centres, mouth corners, brow centres, eyelid
points, eye corners) + timestamps (and event injections) → expected PNG. Plus
fixtures/signals/: face data → expected signal values (6.1). Face data is injected, so no
tracker runs in a golden test.
10.5 [S] §5.3 — Renderers agree "within a tolerance". [P] Proposed tolerance: per
channel |Δ| ≤ 2/255 on ≥ 99.5 % of pixels and ≤ 8/255 on all, after compositing over an
opaque test background.
(Counted under its leading [S]; the tolerance value itself is a proposal.)
10.6 [P] Initial state. At start every layer takes its manifest visible; sprites and
emitters are idle except as 7.6 and 7.7 start them; no clip plays until its autoplay start or an
action (§11.7).
10.7 [P] Particle cap is per layer (maxParticles) per instance; the cross-layer 300
total is a validation rule, not a runtime pool.
10.8 — Withdrawn: a lens has no screen layers, so face triggers and screen layers cannot mix.
10.9 [P] Forward path for fields. Because parsing is strict, a field added after the tag
must come with a field-level engine level (a ENGINE_LEVEL_BY_FIELD table feeding
computeMinEngineVersion), so an older SDK reports "engine too old" instead of "invalid
manifest". No manifest change — a rule for how the editor computes minEngineVersion.
format-v1.1.0 is the first use: ENGINE_LEVEL_BY_FIELD (clip lane, drive, a key's curve,
any presence), ENGINE_LEVEL_BY_ANIMATION_PROPERTY (8, 9) and ENGINE_LEVEL_BY_EASING (8), all 2.
10.10 [P] Host face data (§4.2) must be supplied in the same un-mirrored image space
as the frame (1.2); otherwise left/right anchors swap.
10.11 [P] Observation, no change proposed: there is no mouth anchor and head.tilt is
unsigned, so "tilt left" vs "tilt right" and "from the mouth" effects are not expressible in
v1. Adding an anchor value or event name later is non-breaking (a new engine level). The
full-face anchor (10, 2.2a) was added before anything shipped, so it is engine level 1.
10.12 [P] Rendering is per frame and stateless apart from trigger arm state, sprite
clocks, particle pools with their emission state, PRNG streams, and each layer's playing clip
with its start time and the pending autoplay clip starts (§11.7) — the full list of per-instance
state an SDK keeps.
#11. Keyframe animation
A face lens's optional animations (manifest-v1 "Animations") each animate one layer over their
own time. The value, easing and clip-clock rules are the screen effect's, pinned once.
11.1 [P] Value, easing, clip clock. Exactly effect-rendering-v1.md E10.1 (the value at a
moment), E10.2 (the seven easings) and E10.3 (starts, visibility, looping, end, one clip per
layer), with lens time (8.2) in place of effect time.
11.2 [P] Starts. An autoplay clip starts when lens time reaches its startMs (absent =
0), its clock at startMs — a scheduled start (6.9 step 1a). A play animation action (7.11)
starts or restarts its clip on the frame of the event that fired it (lens.started on the first
frame, a face-signal crossing, a host event), its clock at 0 at t, and cancels that clip's
pending autoplay start. Starting a clip on a layer replaces the one playing there; properties the
new clip does not animate return to rest at once.
11.3 [P] Frame order (6.9). Step 1a starts the scheduled autoplay clips; step 1b ends
non-looping clips with the sprite clocks; play animation runs in steps 2–4 like every action
(later actions win); step 6 draws every layer at its values at t, so a clip started on this frame
shows its clip-time-0 values on this frame.
11.4 [P] What the values change.
- Placement (1.8; stickers, and sprites on a face): the animated offset, scale and rotation replace
the transform's
offset, scale and rotationDeg; the anchor and followRoll apply as
authored (with followRoll the offset stays in the face frame and the face roll is added to the
animated rotation) unless an engine-2 anchor or roll-follow track animates them (11.9, 11.10).
The face frame is re-evaluated every frame (1.6).
- Opacity (5.4): types 2 and 3 multiply by the animated opacity; type 4, either placement, now
multiplies too, resting at 1. Face particles keep opacity 1 for their life.
- Colour look (4.3):
intensity is the animated one.
- Drawn values are clamped: opacity and intensity to [0, 1], scale to ≥ 0 (back and bounce can
overshoot).
- Face particles (2.6, 6.9 step 1c, 7.9): a particle spawns at the anchor plus the offset its
clip gives at the particle's own emission instant
t_k (a burst or an emit: at t).
Catch-up evaluates each t_k against the clips as they stood at t_k — plays from earlier
frames and autoplay starts at or before t_k count, a clip that ends after t_k is still
playing at t_k, this frame's actions come later — while the anchor and W stay this frame's
(6.9 step 1c). A particle never follows later offset changes.
11.5 [P] No face. Clip clocks run on lens time whether or not a face is tracked.
Face-anchored layers are not drawn without a face (2.5) and reappear mid-clip when it returns.
11.6 [P] Visibility. hide does not stop a clip and show does not restart one —
visibility and playback are independent, as for sprites (7.2).
11.7 [P] State. Each layer's playing clip (id, start time) and the pending autoplay starts,
per instance (10.12). Empty at attach — no clip plays until its autoplay start or an action — and
discarded on detach (8.6).
11.8 [P] Custom curve (engine 2). Easing 8 eases the segment by the key's curve
(x1, y1, x2, y2), a CSS cubic-bézier from (0, 0) to (1, 1):
x(s) = 3(1−s)²s·x1 + 3(1−s)s²·x2 + s³, y(s) = 3(1−s)²s·y1 + 3(1−s)s²·y2 + s³. For the segment
fraction u: u ≤ 0 → 0; u ≥ 1 → 1; otherwise solve x(s) = u by exactly 40 bisection
halvings — lo = 0, hi = 1; 40 times mid = (lo + hi) / 2, if x(mid) < u then lo = mid else hi = mid — and the eased fraction is y((lo + hi) / 2), in IEEE-754 double precision, each
polynomial evaluated as written left to right (3·(1−s)·(1−s)·s·p1 + 3·(1−s)·s·s·p2 + s·s·s; other
groupings may differ by a few ulps, inside the 10.5 tolerance). (x1, x2 ∈ [0, 1]
make x non-decreasing, so the bisection is well defined.) Every other easing is E10.2's.
11.9 [P] Anchor track (engine 2, property 8). Keys hold anchor ids. At clip time c the
anchor point is: before the first key or from the last key on, that key's anchor point; between keys
a and b, P(a) + (P(b) − P(a)) × e, where P(n) is anchor n's point on this frame (2.1, 2.2)
and e is key a's easing of the segment fraction (so Hold stays on a until b's time; back and
bounce may overshoot the segment). This point replaces A in 1.8 (stickers, sprites on a face) and
2.6 (face particles); a particle takes the anchor point of its emission instant t_k (11.4), with
this frame's landmarks.
11.10 [P] Roll follow (engine 2, property 9). A placed layer's roll follow f is its track
value, else its resting value followRoll ? 1 : 0; it is clamped to [0, 1] when drawn. With
R(θ) the rotation by θ (positive clockwise on screen, 1.9): P = A + W · R(f·roll) · (offset.x, offset.y)
and the image turns by f·roll + rotation. At f = 1 and f = 0 a renderer uses 1.8's two
formulas as written (face-frame vectors u, v; image axes), so engine-1 placements stay exact.
11.11 [P] Lanes (engine 2). A clip's lane (1–4, absent = 1) splits a layer's playback:
11.2's "replaces the one playing there", the autoplay rule and 11.7's state are per layer and
lane. Clips on other lanes keep playing. A property's drawn value is that of the highest lane whose
playing clip has a track for it; if no playing clip has one, it rests. A non-looping clip that ended
with end behaviour 2 is not playing (its lane falls through to lower lanes); one that ended with 1
holds its last values and keeps its lane.
11.12 [P] Signal-driven clips (engine 2). A clip with drive: { signal } starts like any
clip (autoplay at startMs, or action 6), and from then on is always playing: its clip time on a
frame is clamp01(s) × durationMs, where s is that frame's value of the signal (6.1, after the
preview's overrides), or 0 while no face is tracked. loop and endBehavior have no effect.
Particle catch-up (11.4) evaluates a driven clip with this frame's s at every t_k (the only
signal value the frame has).
Golden scenarios (packages/render/fixtures/scenarios/lens/, hand-checked):
animation-sticker.json (a sticker keyed on a still face: offset, scale, rotation, opacity, back
easing; followRoll on a tilted face), animation-mouth-open.json (a mouth.open crossing starts
a clip), animation-particles.json (spawn offsets sampled at t_k), animation-colour-look.json
(an intensity fade). Engine 2: animation-anchor.json (11.9), animation-roll-follow.json (11.10),
animation-curve.json (11.8), animation-lanes.json (11.11), animation-driven.json (11.12).