A time-based composition model for oxideav: a Scene is a canvas
populated with Objects (images, videos, text, shapes, audio cues)
animated over a timeline. Scenes are the foundation for three distinct
workloads:
- Document layout — a PDF page is a single-frame scene with text, vector shapes, and image objects laid out in their native coordinate system. Edits (adding a watermark, moving an image, rewrapping a paragraph) happen on the scene, not on rasterised pixels, so text stays selectable and vectors stay crisp on re-export.
- Live streaming compositor — a long-running scene fed by external
operations (
AddObject,MoveObject,FadeOut). Intended to sit behind an RTMP server so a remote control plane can drive a per-viewer overlay: add a lower-third during a goal, slide a logo in, trigger a sound effect. - Non-linear video editor (NLE) timeline — Premiere/Resolve-style multi-track editing. Tracks are ordered groups of scene objects, transitions are keyframed cross-fades / wipes, effects are filter chains attached to a single object.
Zero C dependencies — pure Rust, same rules as the rest of oxideav.
Type model complete; vector renderer landed. This crate ships the type model + public-API shape for all three use cases, plus a concrete raster renderer. Encoding and file-format I/O are still follow-ups.
-
Scene,SceneObject,ObjectKind,Transform,Animation,Keyframe,Easing,AudioCuetypes are in place. -
RasterRendereris a concreteSceneRenderer: it walksScene::sampled_at(t)in paint order and composites the vector slice of a scene to an RGBA8VideoFrame. It handles:- backgrounds — solid / transparent / linear + radial gradient, plus
Background::DecodedImage(Arc<VideoFrame>)(a pre-decoded RGBA8 backdrop stretched edge-to-edge); Shapeobjects — rect with corner radius, polygon, SVG path data;ObjectKind::Vectorframes;ObjectKind::Groupcontainers — children resolved by id and inlined under the group's transform / opacity / clip; cycles terminated, missing ids dropped;ObjectKind::Image(ImageSource::Decoded)— the carriedVideoFramewrapped in aNode::Imagespanning its natural size (RGBA8-stride conventionwidth = stride / 4,height = data.len() / stride), sampled through the downstreamoxideav_raster::Renderer's configuredImageFilter;ObjectKind::Video(VideoSource::DecodedFrames { frames, frame_duration })— at scene timetthe renderer picks the in-range frame (finished clips freeze on their final frame;frame_duration <= 0falls back to frame 0).
Decoder-bound
Path/EncodedBytesimage/video/background variants skip silently — pre-decode upstream and feed back via theDecodedvariants.ObjectKind::Live/ObjectKind::Textare skipped pending a font-registry / live-source-aware renderer.Canvas::Vectorscenes are rejected withError::Unsupported(they export theirVectorFramedirectly without rasterisation). - backgrounds — solid / transparent / linear + radial gradient, plus
-
Geometry introspection —
ImageSource::natural_size()/VideoSource::natural_size()/frame_at(t, lifetime_start)(and thecontent_size()accessors that route through them) report natural pixel dimensions and resolve the visible frame for theDecodedvariants; encoded variants returnNone. -
SVG path data —
svg_path::parse_path/parse_svg_path(re-exported at the crate root) lowers an SVG 1.1 path-data string into anoxideav_core::Path, covering the entire grammar (M / L / H / V / C / S / Q / T / A / Z, both cases). Elliptical arcs lower intoPathCommand::ArcTowith the SVG 1.1 F.6.2 out-of-range rules applied at parse time;oxideav-rasterflattens them into cubics downstream. FeedsShape::Pathrendering and bbox queries. -
Audio mixing —
RasterRenderer::render_at(scene, t)mixes the scene'sAudioCues into theRenderedFrame::audioslot, tracking an internalaudio_cursorand emitting a monoVec<f32>covering[audio_cursor, t)atscene.sample_rate. Supported sources:Generator::Silence/SineWave/WhiteNoise(chunk-independent xorshift),PcmS16, andPcmF32. Stereo / multichannel PCM downmixes by averaging; differing source rates resample by nearest-neighbour; the summed mix is multiplied by each cue'svolumeAnimationand clipped to[-1.0, 1.0]. The free functionmix_cues(scene, start, end)exposes the same mixer standalone. Decoder-boundAudioSourcevariants skip silently. -
Traits —
SceneRenderer+SceneSamplerare defined;StubRendererremains the always-Error::Unsupportedplaceholder. -
Paint —
Paint+Gradienttyped paint patterns (multi-stop linear / radial) live in [paint], withBackground::Gradient(_)the richer alternative to the legacy two-colourBackground::LinearGradient. -
Operations DSL —
Scene::apply(op)/apply_batch(ops)drive the [Operation] DSL in-process (add/remove objects, set transforms, animate / cancel, fire audio cues), returning receipts to the caller. -
Compositing helpers —
Scene::merge(other, time_offset, z_offset)splices another scene on (shifting lifetimes / keyframe times, offsetting z-order, extending a finite duration);Scene::next_object_id()allocates a collision-free monotonic id. -
3D-adjacent typed surfaces (landing places for glTF / USD / OBJ importers; the 2D
RasterRendererignores them for now):light—Light/LightCommon/SpotParamspunctual-light primitive (Directional / Point / Spot) per the glTF 2.0 punctual lights extension, withdistance_attenuation/cone_attenuation/irradiance_athelpers.LightInstanceis the pose-carrying wrapper;Scene::lightsis the top-level list (push_light/has_lights/lights_filter).material—Material/PbrMetallicRoughness/AlphaModemetallic-roughness model per the glTF 2.0 core spec, with derived BRDF inputs (diffuse_color/f0/alpha_roughness/fresnel) as methods.Scene::materialsis the palette (push_material/material/materials_filter).node—Mat4/NodeTransform/SceneNode/NodeGraphcolumn-major node-transform graph per the glTF 2.0 core spec. Right-handed,+Yup,+Zforward.Mat4carries the full linear-algebra kit (mul,transpose,determinant,inverse,decompose_trsback to translation/quaternion/scale —Noneon sheared / non-affine / collapsed input) plus quaternion ops (quat_mul/quat_slerpshortest-path spherical interpolation /quat_from_axis_angle/ conjugate / normalize).NodeGraphresolves hierarchy: one-passglobal_matrices, cycle-safevisit/visit_subtree,ancestors/descendants/path_from_root/find_by_name/parent_indices, andvalidate()returning typedNodeGraphErrors (index bounds, multi-parent, cycles, non-finite transforms, zero-length rotations).node_animation— keyframe animation of node TRS properties per the glTF 2.0 animation model (§3.11 + Appendix C):NodeAnimation/AnimationChannel/AnimationSamplerwithStep/Linear(slerp for rotations) /CubicSpline(Hermite with duration-scaled tangents, rotations re-normalised) interpolation, exact-timestamp passthrough, end-clamping,validate(&graph)with typedNodeAnimationErrors, andlocal_transforms_at/global_matrices_atposed resolution.Scene::node_graph+Scene::node_animationscarry both on the scene;Scene::mergeconcatenates lights and materials verbatim and rebases node-graph indices (children, roots, channel targets) so merged hierarchies keep their meaning.
-
No
oxideav-codecor container integration yet — that comes after the render pipeline is real.
pub struct Scene {
pub canvas: Canvas, // pixel dims OR a vector-coord PDF page
pub duration: SceneDuration, // Finite(dur) | Indefinite (streaming)
pub time_base: TimeBase, // rational tick granularity
pub framerate: Rational, // output render cadence (e.g. 30/1, 24000/1001)
pub sample_rate: u32, // audio rate for the mix bus
pub background: Background, // solid colour / image / gradient / transparent
pub objects: Vec<SceneObject>, // z-ordered painter's algorithm
pub audio: Vec<AudioCue>, // triggered by timeline position
pub metadata: Metadata, // author / title / colour-space hints
pub pages: Option<Vec<Page>>, // Some(_) → paged-content mode (PDF / TIFF / EPUB)
pub lights: Vec<LightInstance>, // 3D punctual lights for scenes carrying 3D content
pub materials: Vec<Material>, // PBR material palette for 3D round-trips
pub node_graph: NodeGraph, // 3D node hierarchy (placement)
pub node_animations: Vec<NodeAnimation>, // keyframe TRS animations over node_graph
}A scene is addressed in its own time_base — same rational type oxideav
uses everywhere. framerate is separate: time_base sets the tick
granularity of every scheduled event (keyframe, lifetime, audio cue
trigger); framerate sets the cadence at which the renderer samples the
scene and emits frames to a sink. A scene at time_base = 1/1000 (ms)
and framerate = 30/1 renders at t = 0, 33, 66, 100, … ms. Videos
included via ObjectKind::Video are retimed by the renderer so their
per-frame PTS aligns with this cadence.
SceneDuration::Indefinite signals a streaming scene: no end, no
rewinding, the composition is driven forward by wall-clock time +
operation messages.
pub enum Canvas {
/// Pixel-based raster canvas. NLE + streaming compositor use this.
Raster { width: u32, height: u32, pixel_format: PixelFormat },
/// Unit-agnostic vector canvas. PDF pages use this — the unit is
/// whatever the producer declared (pt, mm, px). All coordinates
/// inside the scene live in this unit; rasterisation happens at
/// export time.
Vector { width: f32, height: f32, unit: LengthUnit },
}Keeping both raster and vector under one type lets the same
SceneObject/Animation/Transform primitives drive PDFs,
compositor streams, and NLE timelines without forking the API.
pub struct SceneObject {
pub id: ObjectId, // stable across edits/operations
pub kind: ObjectKind, // what it IS
pub transform: Transform, // where it is, right now (base state)
pub lifetime: Lifetime, // [start, end) in scene time
pub animations: Vec<Animation>, // per-property keyframe tracks
pub z_order: i32, // painter's algorithm tie-break
pub opacity: f32, // 0.0..=1.0 base opacity
pub blend_mode: BlendMode, // normal, multiply, screen, …
pub effects: Vec<Effect>, // filter chain (blur, colour shift, …)
pub clip: Option<ClipRect>, // geometric clipping region
}pub enum ObjectKind {
/// Static bitmap — PNG/JPEG/raw, decoded upstream into a VideoFrame.
Image(ImageSource),
/// Video stream — consumed as a `Packet` iterator + decoder. The
/// scene's clock drives the stream's PTS; seeking is handled by
/// the underlying demuxer if it supports it.
Video(VideoSource),
/// Styled text run. Preserves font / size / weight / colour
/// metadata so PDF export can emit real text strings and NLE /
/// compositor rasterise through a text-shaping backend.
Text(TextRun),
/// Vector shape — rect, rounded rect, polygon, bezier path.
Shape(Shape),
/// Container object. Applies its own `Transform` before children.
Group(Vec<ObjectId>),
/// Live feed from an external source (RTMP input, camera, etc.).
/// Packets arrive asynchronously; the compositor uses the most
/// recent frame available at render time.
Live(LiveStreamHandle),
}ImageSource / VideoSource / LiveStreamHandle own the heavy
resources — e.g. a VideoSource holds a demuxer + decoder pair, so
copying a SceneObject is cheap but cloning the underlying pixels
requires Arc-shared frame storage (managed by oxideav-core).
pub struct Transform {
pub position: (f32, f32), // canvas units
pub scale: (f32, f32), // 1.0 = natural size
pub rotation: f32, // radians, around anchor
pub anchor: (f32, f32), // 0.0..=1.0 normalised pivot
pub skew: (f32, f32), // radians (Premiere-style)
}
pub struct Animation {
pub property: AnimatedProperty,
pub keyframes: Vec<Keyframe>, // time-sorted
pub easing: Easing, // segment-level default
pub repeat: Repeat, // once / loop / ping-pong
}
pub enum AnimatedProperty {
Position, Scale, Rotation, Opacity, Skew, Anchor,
EffectParam { effect_idx: usize, param: &'static str },
Custom(String), // SceneObjectContent defines semantics
}
pub enum Easing {
Linear, EaseIn, EaseOut, EaseInOut,
CubicBezier(f32, f32, f32, f32), // CSS / AE compatible
Step(usize), // N stepped frames
Hold, // no interpolation — discrete
}Keyframe values are typed per property (Vec2, f32, colour, etc.)
via a KeyframeValue enum that interpolate(a, b, t, easing) acts on.
The transform layer composes upward into per-object and scene-wide AABB accessors so layout, selection, and culling layers don't need to walk the object list by hand:
Shape::content_size()/ObjectKind::content_size()/SceneObject::content_size()report the object-local(width, height)for the kinds that carry one intrinsically (Vector's viewport,Shape::Rect/Shape::PolygonAABB,Live'shint_size). Image, video, text and group returnNone— those extents come from the renderer, not the model.SceneObject::bbox(fallback)returns the object's AABB in canvas space: intrinsic content size when known, the caller-suppliedfallbackotherwise; piped throughTransform::bboxand then intersected withClipRectif the object carries one (zero extent on no overlap so callers can cull).Scene::bbox_at(t, fallback)is the union AABB of every live object at scene timet—Nonefor an empty / fully- dead scene, otherwise the smallest axis-aligned box enclosing every contributing object footprint. Dead and clipped-out objects are skipped so they don't pull the union to their corners. Geometric footprint only — opacity, blend mode, and effect chains are not modelled.Scene::hit_test_at(t, point, fallback)returns theObjectIdof the top-most live object whose AABB containspoint. Painter's-algorithm order: higherz_orderwins, ties broken by later insertion. AABB-only — a rotated rect's AABB contains corners the rect itself does not, so a per-pixel picker layered on top of this remains a follow-up.SceneObject::effective_transform_at(t)/effective_opacity_at(t)compose the object's baseTransformandopacitywith anyPosition/Scale/Rotation/Skew/Anchor/OpacityAnimationtracks evaluated att. Position / rotation / skew add, scale multiplies, anchor replaces, opacity multiplies and clamps to0.0..=1.0.SceneObject::sample_at(t)returns aSamplecarrying(id, z_order, transform, opacity, blend_mode, clip)— the per-frame view a renderer needs without re-running the keyframe evaluator itself.Scene::sampled_at(t)produces oneSampleper live object in paint order (z ascending, ties broken by insertion).
pub struct AudioCue {
pub trigger: TimeStamp, // when playback starts in scene time
pub source: AudioSource, // file / clip / generator
pub volume: Animation, // animated 0.0..=1.0
pub duck: Vec<DuckBus>, // other cues to attenuate while playing
}Audio cues mix into a single output bus per scene. The render pass
produces (VideoFrame, AudioBuffer) at each timestamp; the audio
buffer spans the interval [last_render_time, this_render_time) at
the scene's sample_rate.
Scene + t → SceneSampler.sample_at(t) → RenderedFrame {
video: Option<VideoFrame>, // None for audio-only intervals
audio: AudioBuffer, // always valid, may be silence
operations: Vec<ExportOp>, // e.g. for PDF export: emit text run X
}
A SceneRenderer walks the SceneObject list in z-order, evaluating
transforms + animations at t, clipping against the canvas, and
compositing via the BlendMode. RasterRenderer implements this for
the vector slice today:
use oxideav_scene::{Background, Canvas, RasterRenderer, Scene, SceneRenderer};
let scene = Scene {
canvas: Canvas::raster(1920, 1080),
background: Background::Solid(0x101820FF),
..Scene::default()
};
let mut renderer = RasterRenderer::new();
renderer.prepare(&scene).unwrap();
let frame = renderer.render_at(&scene, 0).unwrap();
// frame.video is Some(VideoFrame) — an RGBA8 plane at the canvas size.The renderer delegates per-object content fetching to each
ObjectKind's own sampler:
A scene acts as a source of rendered frames. Wrap a Scene plus
a SceneRenderer in a RenderedSource and the resulting value
implements SceneSource: one pull() per frame at the scene's
framerate, timestamps auto-advanced by 1 / framerate. Finite
scenes signal end-of-stream by returning None; indefinite scenes
run until externally stopped.
Consumers implement SceneSink — init(&SourceFormat) once, push
per frame, finalise() at end. The helper drive(source, sink) runs
the pull loop:
use oxideav_scene::{drive, RenderedSource, NullSink, StubRenderer, Scene};
let scene = Scene {
framerate: oxideav_core::Rational::new(30, 1),
..Scene::default()
};
let mut src = RenderedSource::new(scene, StubRenderer); // real renderer goes here
let mut sink = NullSink::default();
// drive(&mut src, &mut sink)?; // when the real renderer landsDownstream crates provide the real sinks — an oxideav-scene-encode
sink that pipes frames into an encoder + muxer, an
oxideav-scene-rtmp sink that writes to an RTMP endpoint, a
WindowSink for live preview, etc. Any of these can slot in without
changing the scene or renderer.
Pixel formats get handled transparently in two places:
Inbound (source → scene): a Video / Image / Live object's
source frames can be in any pixel format the decoder produces —
YUV420P, YUV444P, BGRA, RGB24, NV12, whatever. The renderer converts
them to the canvas's pixel format before compositing via
adapt_frame_to_canvas. Writers of per-object samplers call this
once on each pulled frame; canvases that don't declare a raster
format (vector canvases for PDF export) short-circuit the conversion.
Outbound (scene → sink): when a sink expects a pixel format that
differs from the scene's canvas — e.g. a JPEG writer wants RGB24
while the scene composes in YUV420P — wrap the source in
AdaptedSource:
use oxideav_scene::{AdaptedSource, RenderedSource, Scene, StubRenderer};
use oxideav_core::PixelFormat;
let scene = Scene::default(); // canvas: Yuv420P
let src = RenderedSource::new(scene, StubRenderer);
let adapted = AdaptedSource::new(src, PixelFormat::Rgba);
// adapted.format().canvas now reports Rgba; pulled frames are
// transparently converted on the way out.Both paths delegate to oxideav-pixfmt
— the same conversion matrix used across oxideav.
Imagesamplers hold a cached decodedVideoFrame.Videosamplers advance their demuxer/decoder to the requested PTS and return the most recent frame.Textsamplers shape glyphs via a pluggableTextShapertrait (default: a minimal monospace fallback; real layout engines land as separate crates).Shapesamplers rasterise on demand viaoxideav-raster.
Each page becomes a Scene with Canvas::Vector { unit: Pt, width, height } and one SceneObject per glyph run, image, and vector path.
The scene's duration is Finite(1 frame). Edits (redact a region,
drop a watermark, rewrap a column) happen on the scene graph. When the
user re-exports:
- PDF out — the
SceneRendererwalks the tree and emits PDF operators (Tjfor text,Dofor images,f/Sfor vectors), preserving structure. Text remains selectable, hyperlinks survive, bookmarks stay intact. - PNG / JPEG out — the renderer rasterises at a requested DPI.
A daemon holds one Scene per live channel with duration: Indefinite. A control-plane protocol (JSON over WebSocket, say)
surfaces:
{"op": "add_object", "id": "lower-third", "kind": {"image": "...base64..."}, "transform": {...}}
{"op": "animate", "id": "lower-third", "property": "position", "from": [0, 0], "to": [200, 0], "duration_ms": 800, "easing": "ease_out"}
{"op": "remove_object", "id": "lower-third", "delay_ms": 5000}The compositor renders the scene into a VP9/AV1/H.264 encoder fed to an RTMP muxer. Viewers receive a normal stream; the producer only sees the DSL.
Tracks are SceneObject::Group children with a shared z-order band.
Transitions between clips are implemented as opacity / position
animations that overlap two Video objects. Effects are the
effects: Vec<Effect> vector on each object. Scrubbing + preview
works by driving the SceneSampler at arbitrary timestamps; export
renders the entire duration at the target framerate.
src/
├── lib.rs — module exports + Scene / Canvas root types
├── object.rs — SceneObject + ObjectKind + Transform + BlendMode
├── animation.rs — Animation + Keyframe + Easing + interpolation
├── audio.rs — AudioCue + AudioSource
├── audio_mix.rs — mix_cues(): AudioCue list → mono f32 buffer
├── render.rs — SceneRenderer + SceneSampler traits + StubRenderer
├── raster_renderer.rs — RasterRenderer: concrete SceneRenderer (bg + shapes + vector frames → RGBA)
├── raster.rs — rasterize_vector(): VectorFrame → VideoFrame via oxideav-raster
├── text.rs — TextRenderer: TextRun → RGBA via oxideav-scribe + oxideav-raster
├── source.rs — SceneSource + SceneSink + drive() + RenderedSource + NullSink / FnSink
├── adapt.rs — pixel-format adaptation (inbound + outbound, via oxideav-pixfmt)
├── duration.rs — SceneDuration + Lifetime
├── id.rs — ObjectId (stable, editable)
├── ops.rs — Operation enum for the streaming compositor
├── page.rs — Page (paged-content / PDF mode)
├── paint.rs — Paint + Gradient (multi-stop linear / radial)
├── light.rs — Light / LightInstance (3D punctual lights)
├── material.rs — Material / PbrMetallicRoughness / AlphaMode (3D PBR palette)
├── node.rs — Mat4 / NodeTransform / SceneNode / NodeGraph (3D node transform graph + quaternion ops)
├── node_animation.rs — NodeAnimation / AnimationChannel / AnimationSampler (keyframe TRS animation)
└── svg_path.rs — SVG 1.1 path-data string → oxideav_core::Path
Everything is pub and #[non_exhaustive] on public enums so new
variants can land without an SemVer break.
- Not a vector rasteriser. Shape / path rasterisation is delegated
to
oxideav-raster. - Not a text shaper. The
TextShapertrait is pluggable; text layout is delegated tooxideav-scribe. - Not an NLE UI. This crate is the data model + renderer core; the UI is downstream.
- Not a document parser. PDF / SVG ingest land in
oxideav-pdf/oxideav-svg(both pending) and produceScenes.
MIT — same as the rest of oxideav.