Loading...
Loading...
Clear definitions for the terms shaping AI content creation, provenance, and compliance. From workflow reproducibility to regulatory frameworks.
The generative method that produced a 3D asset — matchable provenance for a generated model.
A still produced as an input/reference toward a 3D asset, not a final. A form of still role in still-image work.
Diffusing directly in a 3D latent space (the prior is genuinely 3D, not 2D-projected), yielding cleaner watertight geometry. A form of 3d generation m...
The rule that a 3D-input still should be deliberately un-cinematic — flat light, deep focus, subject square-on and whole, plain keyable background. A ...
Voice turning metallic and prosody degrading over successive extends as the phonetic manifold is lost. Also known as Audio latent drift. A form of aud...
Matching input texture richness to the target aesthetic — maximise pores/grain for cinema, restrain for self-tape/phone footage. A form of conditionin...
An interface design philosophy where API endpoints are built primarily for machine callers (AI agents) rather than exclusively for human users through...
A chronological record of all operations performed by or with an AI system, including inputs, outputs, configuration changes, and user interactions. R...
The practice of transparently communicating to audiences that content was created with or by artificial intelligence. Disclosure requirements vary by ...
A verifiable record of an AI-generated asset's origin, including the model, prompt, parameters, and workflow used to create it. Provenance enables rep...
Under the EU AI Act, a natural or legal person that uses an AI system under its authority, as distinct from the AI provider who develops or markets th...
Structured information embedded in or associated with AI-generated content that describes how the content was produced. This includes model identifier...
A digital asset management system designed from the ground up for AI-generated content. Unlike traditional DAMs adapted for AI workflows, an AI-native...
Read definitionThe continuous background environmental sound bed that establishes place. Also known as Atmos, or Room tone. A form of stem in audio production.
A baked map darkening contact creases and cavities to fake soft contact shadowing. Also known as AO. A form of surfacing in 3D work.
Placing explicit continuity language in the first sentence to suppress the model's default to insert cuts. A form of prompt craft in cinematography.
Photography of buildings and structures emphasising line, scale, and perspective. A form of subject in still-image work.
The width-to-height proportion of the frame (1:1, 3:2, 16:9). A form of composition in still-image work.
The recorded chain from source still → model version → retopo → licence tier for a generated 3D asset. Also known as Lineage. A form of production han...
Dedicating ~14B parameters to the video stream and ~5B to the audio stream, linked by shared timestep conditioning. A form of model architecture in ci...
The fixed total weight (1.0) the attention mechanism distributes across all tokens in a prompt. Also known as Zero-sum attention. A form of prompt cra...
The active navigator that aggregates, weights, and shifts across the entire vector field to determine the most probable next state. Also known as Atte...
Carrying the acoustic state across generated segments without an audible seam or identity break.
The model carrying the acoustic state (room tone, mid-sentence) into a continuation; risks acoustic creep. Also known as Acoustic relay. A form of aud...
A bounded unit of audio content: a cue, a bed, or a transition between sections.
The depth layer a sound occupies in the mix (foreground, midground, background).
Latent-blending audio vectors for a phase-aligned cross-fade that prevents the DC-offset click of a hard cut. Also known as Audio latent blend. A form...
How the model renders speech and sound: computing phonetic timing then translating it to waveforms.
Unified single-pass generation gives excellent per-take sync but ties vocal identity to each individual seed. Also known as Acoustic coupling. A form ...
Using quality signals, behavioral patterns, and visual analysis to surface the highest-value assets from a large generative library without manual rat...
The most distant sonic layer (ambience, walla) sitting behind the mix. A form of audio plane in audio production.
Isolating the subject on a solid or transparent background before upload; busy or subject-coloured backgrounds corrupt reconstruction. Also known as C...
A hard cast shadow in the input sculpted into the mesh as a physical dent; cheapest to prevent, brutal to fix. A form of 3d failure mode in 3D work.
The primary low-resolution DiT stage that builds structure/identity/motion; only ~17% of wall-clock yet does the creative work. A form of sampling and...
Downloading multiple AI-generated images simultaneously as a ZIP archive from a platform like Midjourney. Batch exports introduce specific failure mod...
A sustained underlying layer of music or ambience over which foreground elements sit. Also known as Ambient bed, or Music bed. A form of audio form in...
Decoupled streams constantly exchanging information so visual cues map to auditory events with sub-frame precision. Also known as Cross-modal attentio...
A rough, untextured massing pass that establishes scale and layout before detailing. Also known as Graybox, or Greybox. A form of geometry in 3D work.
Generating extra seconds so the clean middle can bridge a blend, never inheriting tail-end trash pixels. A form of production workflow in cinematograp...
A frontal key placed above the lens, casting a butterfly-shaped shadow under the nose. Also known as Paramount lighting. A form of lighting in still-i...
The Coalition for Content Provenance and Authenticity (C2PA) is an open technical standard for certifying the origin and history of digital content. I...
An unposed, spontaneous capture of a subject. Also known as Reportage. A form of subject in still-image work.
The coordinate (the weighted mean) where the cumulative pull of all prompt tokens converges; the denoising target. Also known as CoG. A form of latent...
System-2 refinement of complex constraints across the latent path using an inference budget, mimicking human deliberation. A form of inference and rea...
The finding that reasoning happens along the latent trajectory, with structural decisions made in specific denoising windows. Also known as CoS. A for...
Forcing calculation of intermediate steps (e.g. how a voice should sound) before rendering the final frame or waveform. Also known as CoT. A form of i...
Strong contrast between light and dark used for dramatic modelling. A form of lighting in still-image work.
The DiT's trained default to introduce dynamic camera moves (pans, tilts, dolly pushes) mimicking cinematic training data. A form of model architectur...
The inference-time overshoot mechanism extrapolating between conditioned and unconditioned predictions to amplify prompt adherence. Also known as CFG....
Framing tight on a subject's face or a detail. Also known as CU. A form of shot type in cinematography.
The human habit of adopting AI outputs with minimal scrutiny, risking override of human intuition and deliberation. A form of inference and reasoning ...
Creating lightweight copies of asset collections that maintain lineage to the original without duplicating underlying files. Analogous to Git branches...
The warmth or coolness of a light source measured in Kelvin. Also known as Kelvin, or White balance. A form of lighting in still-image work.
A node-based visual programming graph in ComfyUI that defines the complete image generation pipeline. Each node represents an operation (model loading...
How prompts and reference latents steer the trajectory: CFG, hard/soft conditioning, identity anchoring, state-continuity.
Hidden, concave, or back-facing geometry the model never saw and invents, usually wrongly. A form of 3d failure mode in 3D work.
A consumer-facing implementation of the C2PA standard that displays a tamper-evident provenance record for digital content. Content Credentials show w...
A storage model that identifies files by a cryptographic hash of their content rather than by file path or name. Two files with identical binary conte...
The 1,000-token prompt memory limit; beyond it, earlier instructions lose their gravity. A form of embedding and tokens in cinematography.
A conditioning model constraining generation to a pose, depth, edge, or scribble map. A form of generation parameter in still-image work.
The automatic grouping of AI-generated assets into discrete creative sessions by detecting boundaries from temporal gaps, parameter changes, and tool ...
Injecting a percentage of one shot's end-vectors into the next shot's high-noise onset to carry identity, lighting, and camera momentum. Also known as...
The challenge of maintaining a continuous provenance chain when creative work flows across multiple AI and traditional tools — from ComfyUI to Photosh...
A discrete bounded segment of music or sound scored to a specific moment or scene. Also known as Music cue. A form of audio form in audio production.
A system for organizing, storing, retrieving, and distributing digital files. Traditional DAM platforms manage photos, videos, and documents with meta...
Read definitionA crisp logo or text reprojected as a texture overlay rather than modelled as geometry. A form of surfacing in 3D work.
Reducing polygon count of a dense mesh while preserving silhouette, for real-time or LOD use. Also known as Decimate. A form of production hand-off in...
Translates the fully denoised latents back into visible pixels and audible sound waves. Also known as VAE decoder. A form of model architecture in cin...
Using modality-specific VAEs to compress signals into separate video and audio latents rather than one shared space. A form of model architecture in c...
The identification and elimination of duplicate assets in a library. Exact deduplication uses content hashing to detect byte-identical files stored un...
Everything-sharp input focus (the opposite of shallow depth of field), preserving the edge cues the reconstructor needs. A form of input-image craft i...
Neon-glowing edges, crushed blacks, and crunchy texture from high CFG multiplying structural error late in a take. A form of sampling and noise schedu...
A structured collection of final AI-generated assets prepared for client handoff, including the image files in required formats, usage rights document...
The small per-step calculation that subtracts a specific amount of randomness, nudging latents toward a likelier coordinate. A form of latent geometry...
The iterative method transforming random noise into a structured audiovisual sequence via per-step delta vectors in latent space. Also known as Denois...
How much the input image is altered (0 = identical, 1 = ignored) in image-to-image generation. Also known as img2img strength. A form of generation pa...
A synthesised or layered non-literal effect created for impact (whoosh, riser, sci-fi tone). Also known as Sound design. A form of sfx kind in audio p...
The recorded or generated spoken lines of on-screen or off-screen characters. Also known as DX, Dialog, VO, Voice-over, or Voiceover. A form of stem i...
Whether a sound originates inside or outside the story world.
Sound whose source exists within the story world and can be heard by the characters. Also known as Source audio. A form of diegesis in audio productio...
Generating structure from chaos by iteratively removing Gaussian noise to reveal an image or sound. A form of sampling and noise schedule in cinematog...
The 22B-parameter asymmetric hybrid marrying transformer scaling with diffusion fidelity; the shot engine's structural backbone. Also known as DiT. A ...
A transformer-backbone diffusion model; a common backbone for 3D generation. Also known as DiT. A form of 3d generation method in 3D work.
A map that physically offsets surface geometry — true depth rather than faked normal detail. Also known as Heightmap. A form of surfacing in 3D work.
Unintended departure of identity/style across frames or versions (the visible symptom of entropy creep). Also known as Latent drift. A form of generat...
The Extend function treating a drifting tail as the new truth, so the next shot starts broken and collapses faster. A form of failure mode in cinemato...
Front-loading high CFG (7-10) during structural onset to lock identity, then decaying it as entropy creeps in. Also known as CFG decay. A form of cond...
A numerical vector representation of words, images, or audio capturing inherent properties and semantic relationships. A form of embedding and tokens ...
The numerical representations the model reasons over: text/latent/thinking tokens, positional and rotary embeddings, the context window.
A high-dimensional mathematical space where images and text are represented as numerical vectors such that semantically similar content occupies nearb...
A measure of unpredictability/disorder in the latent space; generation moves from maximum (noise) to low (structure). A form of sampling and noise sch...
The cumulative slow-burn rise of disorder from compounding FP8/BF16 rounding errors that progressively dissolves the structured manifold, producing dr...
Inability to transition between distinct spatial manifolds (interior/exterior, room/room) in a single take without collapse. A form of failure mode in...
A subject portrayed within their characteristic surroundings to convey context. A form of subject in still-image work.
Opening wide that situates location/time of a scene. A form of wide shot in cinematography.
The European Union's comprehensive regulation governing artificial intelligence systems. Article 12 requires providers of high-risk AI systems to main...
Systemic video-generation failure states and hard walls: text warping, noodle hands, environment lock, drift inheritance, tail-end collapse.
A secondary, softer light reducing shadow contrast from the key. A form of lighting in still-image work.
The last large sigma jump to 0.0 that removes haze and sharpens high-frequency detail (pores, fabric grain, hair edges). A form of sampling and noise ...
Soft, even, shadowless illumination on the input, so cast shadows are not sculpted into the mesh as dents. Also known as Shadowless. A form of input-i...
Luminance mismatch between two stitched shots causing the seam to visibly jump; smoothed by the glue schedule. Also known as Flicker. A form of sampli...
An asset management approach where images are organized primarily by their storage location in a hierarchical folder tree. Folder-based systems force ...
Performed everyday sounds (footsteps, cloth, prop handling) recorded in sync to picture. A form of stem in audio production.
The closest, most prominent sonic layer, usually dialogue or a key effect. A form of audio plane in audio production.
Foreground elements (doorways, arches, windows) that frame the subject. Also known as Frame within a frame. A form of composition in still-image work.
A clean-topology, low-poly, UV'd mesh that drops into an engine with minimal or no retopology (e.g. image-to-3D tool output). Also known as Gameready....
Random uncorrelated values following the normal distribution; the clean-slate raw material diffusion sculpts. A form of sampling and noise schedule in...
A radiance-field scene represented as oriented 3D Gaussians, for real-time novel-view rendering. Also known as 3DGS, Splat, or Splatting. A form of ge...
The instruction-tuned text encoder (3,840 hidden dims, 262,208 vocab) parsing prompts into the shared latent manifold. Also known as Gemma 3 12B, or T...
The chain of parent-child relationships between AI-generated assets — from an initial generation through variations, upscales, and refinements. Unlike...
Latent-generation control surface that materially affects matchable output.
Initial latents treated as ground-truth DNA dictating how identity and lighting evolve across the entire sequence. A form of conditioning and guidance...
How a 3D asset's form is represented: polygon mesh, point cloud, Gaussian splat, or rough blockout.
The high-sigma early steps that decide global geometry and bone structure rather than colour or texture. Also known as Rough sketch phase. A form of s...
A low-noise refinement schedule (0.25, 0.18, 0.12, 0.06, 0.0) that smooths a blend seam without changing geometry. Also known as Transition sigma sche...
A compositional proportion (~1.618) placing key elements along a spiral or division for natural balance. Also known as Golden spiral. A form of compos...
Using brackets for global tone and em-dashes to scope cues to specific tokens, preventing attention bleed. A form of prompt craft in cinematography.
The attractive pull tokens exert on the trajectory; early prompt content weights more heavily, so anchors go first. Also known as Front-loading. A for...
Over-amplifying low-frequency signals via high CFG, producing deep-fried colours and unrealistic contrast. A form of sampling and noise schedule in ci...
Hair and fur reconstructed as solid bulk mass rather than strands or cards; hero work grooms hair natively instead. A form of 3d failure mode in 3D wo...
A specific sync-to-picture effect tied to a visible on-screen action (door slam, gunshot). Also known as Spot effect. A form of sfx kind in audio prod...
The starting frame compressed into latents that act as the literal physical denoising state the model extends. Also known as Hard-conditioning latents...
The space between the subject's head and the top edge of the frame. A form of composition in still-image work.
A final, presentation-grade single image. A form of still role in still-image work.
Tightly clustered early sigma steps (1.0 to 0.975) spending ~50% of compute to lock identity and fight identity popping. Also known as High-noise end....
A dense, high-detail mesh — a digital sculpt or the source before decimation. Also known as Highpoly, or Sculpt. A form of mesh in 3D work.
The model filling openings and cavities that should stay hollow, producing solid where the subject was open. A form of 3d failure mode in 3D work.
Treating the video-generation model as one stage among image-generation, TTS, NLE/VFX, and audio-post tools. Also known as Multi-tool stack. A form of...
A retrieval architecture that combines structured metadata queries (exact filters on tool, date, model) with vector similarity search (semantic meanin...
The t=0 frame that establishes and holds subject identity; it degrades as the model applies its transition function over time. A form of conditioning ...
Characters resolving as different people when insufficient high-noise compute fails to lock a stable identity manifold. Also known as Identity pop. A ...
Generating a 3D asset from a single reference image. A form of 3d generation method in 3D work.
How the model deliberates: System-1 vs System-2 processing, thinking tokens, chains of latents/steps/thought, statistical inference.
The compute allotment a System-2 pass spends deliberating over prompt constraints before diffusion begins. A form of inference and reasoning in cinema...
The multi-stage processing system that transforms a raw uploaded file into a fully indexed, searchable asset. Stages include content hashing for dedup...
Regenerating a masked region of an image while preserving the rest. Also known as Masked generation. A form of generation parameter in still-image wor...
How to frame the source still for image-to-3D: the reconstructor imagines geometry from pixels, so ambiguity becomes hallucinated geometry.
Adding descriptive text reduces adherence to primary intent because every word competes for finite attention weight. A form of prompt craft in cinemat...
The 2025.1 revision of the IPTC Photo Metadata Standard that introduces dedicated fields for AI-generated content. New fields include AISystemUsed, AI...
The primary, dominant light source shaping the subject. A form of lighting in still-image work.
Pruning to the load-bearing tokens; adherence is ~80% at 3 instructions, ~50% at 8, under 28% at 15. Also known as KISS. A form of prompt craft in cin...
Using the minimum words to convey a command, keeping the weighted mean a sharp point rather than a blurry cloud. Also known as Laconic precision. A fo...
A feed-forward model that reconstructs 3D geometry from one or a few images in seconds. Also known as LRM. A form of 3d generation method in 3D work.
Mixing the generated latents of two shots inside diffusion (e.g. at sigma 0.5) so features settle on a shared mathematical middle ground. Also known a...
The geometry of the latent space: manifolds, vectors, trajectories, and the forces that pull generation toward a target.
The high-dimensional space in which the model represents data by essential semantic features rather than raw pixels; similar concepts sit close togeth...
A VAE-produced token representing a patch's textures/colours/edges; the raw diffusion material. Also known as Visual token. A form of embedding and to...
The continuous mathematical path the initial latents follow as they transform into generated latents across denoising. Also known as Singular trajecto...
The quantified essence of creative intent — simultaneously a point (state) and an arrow (direction/transformation) in the manifold. Also known as Nume...
Space left in front of a subject's gaze or direction of motion. Also known as Nose room. A form of composition in still-image work.
A set of progressively simpler mesh versions swapped by distance to keep a real-time scene performant. Also known as LOD. A form of production hand-of...
Small-source hard light gives sharp shadows; large-source soft light gives gradual shadows. Also known as Hard light, or Soft light. A form of lightin...
The 2-4 essential tokens — the unavoidable truth of the shot — that anchor the scene's centre of gravity. Also known as Massive anchor. A form of prom...
A key slightly off-axis casting a small loop-shaped shadow from the nose. A form of lighting in still-image work.
A technique for fine-tuning large AI image generation models on a small dataset to learn a specific style, subject, or concept. LoRA files modify a ba...
A mesh with a deliberately small polygon count, for real-time or stylised use. Also known as Lowpoly. A form of mesh in 3D work.
Extreme close-range photography rendering small subjects at life-size or greater. A form of subject in still-image work.
The topological surface holding the set of all mathematically valid variations of a given subject or environment. Also known as Latent manifold. A for...
Polygonal surface representation made of vertices, edges, and faces. Also known as Polygons, or Polymesh. A form of geometry in 3D work.
The architectural shift where AI-generated assets arrive with rich generation metadata already embedded by the creation tool, inverting the traditiona...
A three-stage processing layer that translates tool-specific metadata formats from different AI tools (ComfyUI JSON, Midjourney Discord strings, DALL-...
The process of reconnecting AI-generated images that have lost their generation metadata to their original prompts, parameters, and provenance records...
The removal of embedded metadata from image files before distribution. Creates a tension for AI-generated content: stripping protects creative process...
The intermediate sonic layer sitting between foreground performance and background ambience. A form of audio plane in audio production.
The model defaulting to the most generic average interpretation (e.g. 'person walking') when contradictory cues cannot resolve. A form of conditioning...
Final-stage CFG amplifying within-mode contraction to lock in fine details and reduce variability. A form of conditioning and guidance in cinematograp...
A modern video-generation stack: the diffusion transformer, attention, the VAE, the text encoder, and the dual video/audio streams.
An open protocol that allows AI agents and language models to interact with external tools and services through a standardized interface. In creative ...
Chaotic, rapidly occluding hand topologies the 3D temporal RoPE fails to track, so digits merge or multiply. Also known as Noodle hands. A form of fai...
Generating several view-consistent 2D images with a diffusion model, then reconstructing 3D from them; risks view disagreement / ghosting. A form of 3...
A specialised guider splitting CFG into independent text-guidance strength and cross-modal alignment strength. Also known as Multimodal guidance. A fo...
Treating pixels, audio, and text as numerical vectors in one hidden space, enabling frame-accurate lip-sync. Also known as Multimodal fusion. A form o...
Generating a 3D asset from several overlapping views of one subject; more depth cues than a single image. Also known as Multiview. A form of 3d genera...
Composed musical material scoring a scene, either diegetic or non-diegetic. Also known as Score. A form of stem in audio production.
Text describing what to exclude, steering generation away from unwanted features. A form of generation parameter in still-image work.
An implicit radiance field learned from posed images for novel-view synthesis. Also known as NeRF. A form of 3d generation method in 3D work.
Adding timestep-dependent noise directly to the hard-conditioning latents as the raw material the model refines. A form of sampling and noise schedule...
Sound outside the story world (score, narration) that the characters cannot hear. A form of diegesis in audio production.
Edges shared by more than two faces or otherwise unclean topology that breaks downstream tools until repaired. Also known as Non-manifold. A form of 3...
A tangent-space map that fakes high-frequency surface detail on a low-poly mesh. Also known as Normalmap, or Normals. A form of surfacing in 3D work.
Multiple objects in one input merging into a single fused mass; avoided by one-subject-per-run. A form of 3d failure mode in 3D work.
The Universal Scene Description interchange standard; layering and metadata carry an asset's lineage as first-class data. Also known as USD, or USDZ. ...
A consistent front/side/back/top image set that lets the model recover otherwise-hidden geometry. Also known as Turnaround. A form of input-image craf...
A dense, evenly-triangulated sculpt with no clean edge flow; the reason hero output must be retopologised. A form of 3d failure mode in 3D work.
The VAE dividing a frame into a grid of patches, each compressed into a latent token of textures, colours, and edges. Also known as Patches. A form of...
Physically based materials (albedo / metallic / roughness / normal) that light consistently across scenes. Also known as PBR. A form of surfacing in 3...
The natural rhythm/timing of speech the model computes via thinking tokens before rendering the waveform. Also known as Prosody. A form of audio synth...
Reconstructing 3D geometry and texture from overlapping photographs. Also known as Photoscan. A form of 3d generation method in 3D work.
An unstructured set of 3D points, typically from scanning, prior to meshing. Also known as Pointcloud. A form of geometry in 3D work.
The listening perspective the mix adopts — whose ears we hear from; the audio analogue of point of view. Also known as POA.
The target polygon / face count an asset class is allowed, set by its role (hero vs background prop). Also known as Facelimit. A form of production ha...
The process of progressively filtering a large generative asset library down to a curated portfolio — typically the top 2% of output representing the ...
Encoding token order/position, since transformers are otherwise position-agnostic about sequence. A form of embedding and tokens in cinematography.
A light source visible within the frame (lamp, window) that motivates the scene's lighting. Also known as Motivated light. A form of lighting in still...
A controlled studio capture of a product as the hero subject. Also known as Packshot. A form of subject in still-image work.
Turning a raw generated sculpt into a production-usable asset: retopology, UVs, PBR, interchange, and provenance.
Production techniques and pipeline stages: extend, stitch-and-denoiser, seed variance, quantisation, tiered architecture, upscalers.
The discipline of writing token-efficient prompts: managing the attention budget, instruction dilution, prosody, scoping, anti-cut language.
Changing prompt wording, which fundamentally changes the manifold; the fix for blocking, framing, or grammar problems. A form of production workflow i...
A structured, searchable collection of generation prompts with their parameters, output examples, and performance history. Unlike simple prompt lists ...
The text prompts, negative prompts, and associated generation settings captured alongside an AI-generated asset. Prompt metadata enables search, categ...
A break in the chain of generation metadata that occurs when an AI-generated asset crosses a tool boundary — for example, exporting from Midjourney to...
Typographical symbols inside quotes act as literal acoustic tuning: ! shouts, ? lifts, ellipsis decays, dash pauses. A form of prompt craft in cinemat...
Rebuilding a triangulated sculpt as an even quad mesh with better edge flow for deformation. Also known as Quadremesh, or Remesh. A form of production...
Precision tiers (FP16, FP8, GGUF Q8/Q4) trading VRAM for softening that prompts counter with texture cues. Also known as Quantization. A form of produ...
The set of mathematically valid variations of a subject's identity established from the initial frame. A form of latent geometry in cinematography.
High-dimensional statistical inference (geometric optimisation over embeddings), not conscious cognitive judgment. Also known as Latent logic. A form ...
A flow-matching generative formulation (e.g. a rectified-flow 3D generator) that learns near-straight transport paths for fast, stable 3D shape genera...
A portrait key creating a small triangle of light on the shadowed cheek. A form of lighting in still-image work.
Rebuilding a dense or scanned mesh as clean, animation-ready topology. Also known as Retopo. A form of topology in 3D work.
The skeletal / control structure that lets a 3D model be posed or animated. Also known as Armature, Rigged, or Skeleton.
An A-pose or T-pose with limbs separated from the torso, so the mesh does not fuse arm-to-body and can be rigged. Also known as A-pose, or T-pose. A f...
Rotates embedding vectors in a complex plane by an angle set by token position, encoding relative distance. Also known as RoPE. A form of embedding an...
The numerical solver that denoises the diffusion latent at each step. Also known as Sampling method. A form of generation parameter in still-image wor...
The diffusion denoising process and the sigma/noise schedule that allocates compute across the noise range.
The number of denoising iterations; more steps refine detail at higher compute cost. Also known as Steps. A form of generation parameter in still-imag...
California Senate Bill 942, effective January 2026, requiring providers of generative AI systems to offer provenance tools and users of covered AI sys...
Generated meshes carry no inherent real-world scale or up-axis; both must be standardised on ingest. A form of 3d failure mode in 3D work.
Transformers handling far more parameters than U-Nets without instability, enabling the asymmetric 22B engine. A form of model architecture in cinemat...
The function setting the noise level at each step. Also known as Karras. A form of generation parameter in still-image work.
RNG seed; reproducibility anchor and a strong Signal-1 match key. A form of generation mechanic in cinematography.
Varying the random seed (4-6 takes) to explore different trajectories within the same manifold when performance nuance is off. Also known as Seed rota...
Abstract directions in the vector space where each position corresponds to a feature such as warmth, vertical motion, or metallic texture. Also known ...
Using high-vector-proximity synonyms (e.g. 'parched' vs 'thirsty') to steer texture toward a neighbourhood without more words. A form of prompt craft ...
Synonyms and filler that each consume a token and dilute the gravity of the primary instruction. Also known as Descriptive noise. A form of prompt cra...
A retrieval method that finds content based on conceptual meaning rather than exact keyword matches, by comparing vector embeddings in high-dimensiona...
Steering the denoising trajectory toward the prompt's weighted mean using mathematical anchors rather than literal understanding. A form of conditioni...
Whether an effect is a literal sync-to-picture hard effect or a synthesised designed effect.
The use of AI tools by employees without organizational knowledge, approval, or governance. Similar to shadow IT, shadow AI creates compliance risks w...
Each transformer block running self-attention, text cross-attention, audio-visual cross-attention, and a feed-forward network in sequence. A form of m...
The noise level at a given step; 1.0 is total Gaussian noise and 0.0 is a fully denoised image or audio track. Also known as Noise level. A form of sa...
A non-linear sequence of sigmas that reshapes where in the noise range the model spends its compute budget. Also known as ManualSigmas, or Noise sched...
The deliberate absence of sound used for dramatic or rhythmic effect (cf. the field manual's 'allowable silence' in the speech budget). A form of stem...
Using an input merely as an inspirational visual reference rather than a literal starting state (contrast hard-conditioning). Also known as Visual ref...
The MovieLabs 2030 direction where an asset is a metadata-rich, provenance-carrying package rather than a bare mesh file. A form of production hand-of...
Non-speech, non-music sounds representing the actions and events in a scene. Also known as SFX. A form of stem in audio production.
The VAE's 32x downsampling of physical geometry, why micro-structures like text and hands struggle to survive. Also known as Spatial downsampling. A f...
The heaviest stage (~37% of wall-clock) lifting high-frequency texture from the base sampler's low-resolution foundation. A form of production workflo...
A key at ~90 degrees lighting exactly half the face while the other half stays in shadow. A form of lighting in still-image work.
Overlaid-subtitle hallucination triggered by mentioning dialogue, captions, or text in the prompt. Also known as Subtitle hallucination. A form of fai...
The architectural rule that the model extends states rather than transforms them, carrying the initial manifold forward. Also known as Inertial force....
Calculating the most probable visual and acoustic patterns for semantic vectors from the training distribution. A form of inference and reasoning in c...
A grouped audio mixdown of one category (dialogue, music, or effects) kept separate for mixing and delivery. Also known as Submix.
Generating two shots and blending their overlap with a light denoise pass to purge drift across the seam. Also known as Manual latent blending. A form...
Gradual visual inconsistency that emerges when multiple team members generate AI images without shared style governance. Style drift occurs when creat...
Midjourney's --sref parameter applies a consistent visual aesthetic from a reference image or saved style code to new generations. Style references en...
One subject per generation; multi-object inputs fuse, so separate runs are reassembled in a DCC. A form of input-image craft in 3D work.
How a 3D surface looks: UV layout, physically based materials, textures, and detail maps.
Fast, pattern-based output relying on immediate statistical likelihood; the default mode for most takes. Also known as System 1. A form of inference a...
Uses an inference budget and thinking tokens to evaluate complex prompt constraints before the diffusion pass. Also known as System 2. A form of infer...
Exponential autoregressive failure at a take's end where limbs morph and identity pops into a new person. Also known as Structural collapse. A form of...
Compressing video across time (8x stride) so one latent token packs several frames; the reason on-screen typography is destroyed. Also known as Tempor...
Generalises RoPE to 3D spatiotemporal tokens (t,x,y); low-frequency temporal channels avoid phase wrapping over long clips. Also known as 3D RoPE. A f...
Querying an asset library using time-based expressions — "what I made last Tuesday," "images from the brutalist architecture session," or "everything ...
8x temporal compression loses high-frequency edge data, so on-screen letters drift or morph into illegible glyphs. Also known as Text warping. A form ...
Stabilising micro-motion and frame rate after spatial upscaling. A form of production workflow in cinematography.
Gemma-3-encoded vectors capturing prompt semantics that guide the video and audio generation. A form of embedding and tokens in cinematography.
On-surface text and logos smearing into illegible geometry; treated as a reprojected decal, not modelled as shape. A form of 3d failure mode in 3D wor...
A sub-word unit from the 262,208-token vocabulary; common words are one token, rare words split into several. A form of embedding and tokens in cinema...
Generating a 3D asset from a text prompt. A form of 3d generation method in 3D work.
A learned embedding token capturing a concept or style, invoked by keyword. Also known as TI. A form of generation parameter in still-image work.
An image map applied to a surface — the base colour / albedo and its companion maps. Also known as Albedo, or Basecolor. A form of surfacing in 3D wor...
Transferring high-poly detail (normals / AO / curvature) into maps for a low-poly target. Also known as Bake, or Baking. A form of surfacing in 3D wor...
The large middle sigma jumps that fast-travel through the manifold once structure is set; over-sampling here breeds drift. A form of sampling and nois...
The absence of a common metadata standard across AI generation tools. Each tool uses its own storage location, format, field names, and encoding — Com...
Over-loud delivery on quiet dialogue caused by a stray exclamation mark or capitals inside the quote. A form of failure mode in cinematography.
Strategic allocation of compute via thinking tokens to resolve ambiguity, simulating trajectories before committing. A form of inference and reasoning...
Hidden tokens the encoder generates to deliberate over a prompt before generation, critical for phonetic timing. A form of embedding and tokens in cin...
The standard key + fill + back/rim lighting arrangement. A form of lighting in still-image work.
A front-plus-one-side hero view that yields the most depth from a single input image. A form of input-image craft in 3D work.
Three workflow variants differing in upscaler steps — survey (Rapid), comparison (Fast), shipping (Final). A form of production workflow in cinematogr...
Calibrating dialogue to 2.5-3.0 words per second so the speech fits the shot's duration. Also known as Words per second. A form of prompt craft in cin...
The attention weight/gravity a token carries; load-bearing nouns and verbs are heavy, connector words near-zero. Also known as High-mass token. A form...
The edge/face layout of a mesh; clean topology deforms and textures predictably. Also known as Edgeflow. A form of geometry in 3D work.
An IPTC Digital Source Type standard value indicating that content was generated by an AI model trained on data. Midjourney embeds this value in downl...
Point splines drawn on the initial frame infused into latent space to force generated latents along a precise physical path. Also known as Motion trac...
The attention-based architecture processing all tokens simultaneously for global context, scaling hardware-friendly with more parameters. A form of mo...
A sonic bridge (crossfade, segue, stinger) connecting two sections. Also known as Crossfade, or Segue. A form of audio form in audio production.
Glass, liquid, and other refractive surfaces sculpted as solid opaque lumps rather than thin-walled transparent shells. A form of 3d failure mode in 3...
A compact 3D representation storing features on three orthogonal feature planes; the output of many large reconstruction models. Also known as Tri-pla...
A model that increases resolution and adds high-frequency detail. Also known as Hires fix. A form of generation parameter in still-image work.
The 2D parameterisation that maps textures onto a 3D surface. Also known as UV, UVs, or Unwrap. A form of surfacing in 3D work.
The branching history of how an AI-generated image evolved through successive variations, upscales, and remixes from its original generation grid. Lin...
Compresses raw pixels/audio into latent tokens and decodes them back, using 32x spatial and 8x temporal downsampling. Also known as VAE. A form of mod...
Reasoning by computing the mathematical distance between instruction tokens and latent trajectories in hidden space. Also known as Semantic neighbourh...
Navigating meaning by moving along interpretable latent directions (age, lighting, velocity), inducing semantically coherent transformations. Also kno...
The mathematical distance moved between latent points; drastic state transitions need large leaps the model suppresses for coherence. A form of latent...
The length of a displacement vector, representing the scale of the requested change. A form of latent geometry in cinematography.
One-click extension setting the last frame as new hard-conditioning latents; fast but inherits and compounds drift. Also known as Extend. A form of pr...
Darkening (or lightening) toward the frame edges that draws the eye inward. A form of composition in still-image work.
The component matching typographical prosody triggers (! ? ... dash) to specific frequency and amplitude waveforms. A form of audio synthesis in audio...
Vocal identity is generated fresh per seed/prompt and varies across takes — an architectural, not configurational, limit. Also known as Vocal identity...
Indistinct background crowd murmur or chatter suggesting a populated space. Also known as Crowd. A form of stem in audio production.
A closed, manifold mesh with no holes — required for boolean ops, 3D printing, and clean simulation. Also known as Watertight. A form of production ha...
The model's most probable statistical interpretation of the whole prompt, used as the navigational target during early denoising. A form of latent geo...
The counter-intuitive truth that base sampler is not the wall-clock bottleneck; upscale and decode dominate — profile first. A form of failure mode in...
The ability to recreate an AI-generated asset by replaying its original workflow with identical parameters. Reproducibility requires preserving the co...
Numonic captures the prompts, models, and parameters behind your AI-generated work automatically — so everything in this glossary becomes something you can search.