Making and shipping

DriftCapture

200 symbols, imported from @driftengine/capture.

Explained in DriftCapture.

Classes

Sam21Tracker

Interfaces

BlockNormals
The normal equations in blocks: cameras, points of three, and one link per observation.
BrowserClip
A browser's clip, and how to let it go.
CaptureFrame
CaptureScene
ClipTokenizer
CLIP's tokenizer — byte-level BPE — as Transformers' CLIPTokenizer runs it at the revision the manifest pins, for OWLv2's text queries.
ClipTowerConfig
CollisionOptions
CollisionSource
A mesh as meshShape takes it, and what had to go for it to be one.
DecimateOptions
DelightOptions
DelightOut
What comes back, one entry per vertex of the mesh handed in.
DelightView
One view of the surface: a linear frame and where it was taken from.
DepthAnything2Config
DepthAnything3Config
DepthEstimate
DepthEstimator
DepthView
Descriptors
The descriptions of a frame's features: BYTES bytes each, bit by comparison.
EntityProposal
What a scene could make of a region.
FeatureSet
The features of one frame, in the arrays a caller owns.
FitOptions
FrameChoice
FrameSize
A clip's frames, and which of them a capture keeps.
FrameSource
FuseOptions
GaussianGradients
Where the gradients go: one entry per parameter of every Gaussian.
GaussianSet
Where a Gaussian lands on screen: the projection @driftengine/splats' shader performs, in JavaScript, and the order a frame draws them in.
HeadConfig
HieraConfig
LeastSquares
A least-squares problem: how many residuals, how many parameters, and both at a point.
Light
LitScene
MarchOptions
MaskView
Where a model is prompted, and what came back for each prompt.
MeshHit
MobileSamConfig
OptimiseOptions
Owlv2Config
Owlv2Detection
Owlv2Patches
PoseOptions
PoseResult
PreparedFrame
Projected
One Gaussian as the screen sees it.
PropHulls
A prop's hulls, and how much empty space they added.
ProposeOptions
RasterCamera
RawFrame
A frame as a host hands it over: RGBA bytes, and the size they are.
Region
What a region is, once something has decided where it is.
RelativePoseOut
RelativePoseResult
Sam21Config
Sam21Frame
A frame's result: every object's mask logits at a quarter of the encoder's size, and its score.
Sam21MemoryConfig
Sam21Video
Sam21Weights
The weights each graph reads: a converted file's graphs, or a checkpoint for all four.
SamDecoderConfig
SamPrompt
SamPromptLayout
Where a checkpoint keeps its prompt encoder, and how its processor places a prompt: SAM's and MobileSAM's, or Transformers' SAM 2.
SegmentOptions
SurfaceView
One view's opinion: its depth, how well covered each pixel was, and where it stood.
TestBox
TestCamera
Analytic scenes rendered on the CPU, so a capture's stages can be held to geometry that is known before anything runs.
TestPlane
TestScene
TinyVitConfig
VitConfig
Volume
A field of samples over a box of the world, x fastest, then y, then z.

Functions

accumulateGradients
The gradients of a render against dPixels — the loss's derivative by each rendered channel, four a pixel — added into out. The render is repeated here rather than taken as an argument, because the backward walk needs what each Gaussian contributed on the way.
areaResize
src, width × height pixels of channels bytes, into out at outWidth × outHeight. Only shrinking is defined, which is what a frame's preparation asks for.
browserFrameSource
The clip's frames, addressed at the clip's rate. frameRate is the fallback where the browser cannot say what the rate is.
captureFile
Every stage of scene as one .drft.
cholesky
The lower triangle l with l · lᵀ = m, or false where m is not positive definite — which is how a caller learns that its normal equations are singular, rather than by finding a NaN later.
choleskySolve
Solves l · lᵀ · x = b in place over x, given the Cholesky factor.
clipTokenizer
A tokenizer over merges, [count, 2] token ids in rank order.
collisionMesh
mesh's triangles, less the ones with no area, renumbered onto the vertices that remain.
createDelightOut
createDepthEstimator
An estimator over one model's weights, running its graphs through run.
createFeatureSet
createGradients
createVolume
decimate
mesh simplified until it holds at most budget triangles.
decodeCamera
One view's camera from its pose encoding: worldToCamera, 3 × 4 row-major — the inverse of the camera-to-world rotation and translation the decoder predicts — and intrinsics, 3 × 3 row-major, with focal lengths from the two fields of view and the principal point at the centre.
decodeDepth
Depth and confidence from one view's logits, [2, height, width]: exp of the first channel, and one more than exp of the second, into the caller's maps.
delight
mesh's material, as seen in views.
depthAnything2
The graph for one image of height × width pixels, each a multiple of the patch, normalised by ImageNet's mean and deviation: input image, [3, height, width]; output depth, [1, height, width].
depthAnything3
The graph for views images of height × width pixels, each a multiple of 14, normalised by ImageNet's mean and deviation. Inputs are image0 … as [3, height, width]; outputs are logits0 … as [2, height, width] and pose as [views, 9]. The first view is the reference.
describeFeatures
A description of every feature in features, 32 bytes each.
detectFeatures
The strongest corners of image, RGBA at width × height, into out — at most budget of them, strongest first.
eachFrame
Every frame in order, each read over the last into one buffer the visitor must not keep.
estimatePoses
The clip's camera path, into out: 3 × 4 world-to-camera a frame.
fuseDepth
Every view's surface written into volume, which accumulates rather than replaces.
imageLoss
The squared difference between a render and a frame, and the render's own gradient.
interpolate
A per-vertex attribute at a hit, by its barycentric weights.
labelMasks
A name for each mask index, from a detector's boxes.
levenbergMarquardt
Levenberg–Marquardt over x, in place: the Gauss–Newton step with damping · diag(JᵀJ) added, the damping raised where a step would cost more and lowered where it pays, which is the whole reason it is not Gauss–Newton — a step that overshoots is rejected rather than taken, and a start far from the answer is where that decides between converging and diverging.
liftMasks
Masks from several views lifted onto the mesh, by what most of the views that saw a triangle say.
longestSideSize
The size a frame is resized to, its longest side at size, rounded up at a half.
lookAt
A world-to-camera for an eye looking at a target, with the world's up.
marchVolume
volume's surface, welded so that neighbours share vertices exactly.
matchFeatures
The matches between two frames' descriptions into out, two indices each, and how many there are. A match is mutual and clears Lowe's ratio, which defaults to 0.8.
miniatureCheckpoint
Every tensor the upstream's backbone, head and camera decoder hold, named as the checkpoints are.
miniatureCheckpoint2
Every tensor Transformers' Depth Anything holds, in the older names its checkpoints keep.
miniatureImages
Seeded images for views views, normalised as the model expects, [3, height, width] each.
miniatureOwlv2Checkpoint
Every tensor the upstream's Owlv2ForObjectDetection holds, named as its checkpoints are.
miniatureSam21Checkpoint
Every tensor the upstream's Sam2VideoModel holds, named as its checkpoints are.
miniatureSamCheckpoint
Every tensor the upstream's encoder, prompt encoder and mask decoder hold.
mobileSamDecoder
embedding and prompt, [tokens, dim] — absent when tokens is zero — and mask when refine is set, to masks, [4, 4·grid, 4·grid], and quality, [1, 4].
mobileSamEncoder
image, [3, size, size], to embedding, [neck, grid, grid].
nearestHit
The nearest triangle a ray meets, by Möller–Trumbore over every one of them.
optimiseGaussians
A clip fitted to Gaussians. Nothing here is the caller's array; everything is freshly cut.
owlv2BoxBias
[x, y, w, h] for every cell, row by row: the logits of its far corner and of the cell's size.
owlv2Detections
The detections above threshold among logits, [cells, count], with boxes, [cells, 4] as the image graph answers them, for an original image of height by width.
owlv2Image
image, [3, size, size], to five values a patch: classes, [cells, text.dim], the class embedding before it is normalised; shift and scale, [cells, 1], the logit's shift and its scale before the ELU; boxes, [cells, 4], centre and size in the image's fractions; and objectness, [cells, 1], a logit.
owlv2Logits
out, [cells, count]: every patch's logit for every one of count queries, [count, width] as the text graph answers them.
owlv2Text
tokens, [queries · positions], and ends, [queries], as owlv2Tokens makes them, to queries, [queries, text.dim]: each query's embedding before it is normalised.
owlv2Tokens
OWLv2's text queries as its text graph takes them: tokens, [queries · positions], each query padded with zeros to positions, and ends, [queries], the row of each query's end of text in that list — the upstream pools the row of the largest token, which is that one. A query longer than positions is refused, as the upstream's position table refuses it.
prepareDepthFrame
rgba, width × height pixels of four bytes, as Depth Anything 3 takes it at size, in whole patches of patch.
prepareOwlv2Frame
rgba as OWLv2's processor prepares it: [3, square, square] at the model's size.
prepareSam2Frame
rgba as SAM 2's processor prepares it: [3, square, square], each axis scaled on its own, and normalised by mean and deviation in 0–1.
prepareSamFrame
rgba as MobileSAM's predictor prepares it: [3, square, square], the frame itself in the top left at the size this answers, and zeros after it.
projectGaussian
One Gaussian projected, or null where it is behind the camera or too flat to draw.
promptGrid
A lattice of prompt points over an image, seeded.
propHulls
mesh decomposed into convex hulls the container can hold.
proposalScene
The proposals as a scene, keyed by field id — the form serializeWorld writes.
proposeEntities
Each region as a proposal.
rasteriseGaussians
set through camera into out, four floats a pixel: linear colour, premultiplied as it composites, and the alpha the cloud covered the pixel with.
readProposals
The proposals a scene carries, in the order they were written.
relativePose
The pose of the second view relative to the first, from count correspondences in pixels. random is the caller's seeded generator: a capture answers the same twice.
renderDepth
The cloud's depth and coverage, one of each per pixel.
renderLit
scene seen from camera, into out as linear premultiplied RGBA.
renderTestScene
scene through camera into out, RGBA, and the distance along the camera's own axis into depth where one is given — zero where the ray met nothing.
sam21Decoder
embedding, high1, high0 and prompt, [tokens, dim] — absent when tokens is zero — and mask when refine is set, to masks, quality, object and pointers. With noMemory the decoder adds no_memory_embedding to the features itself, as an image with nothing remembered is decoded; otherwise the features are the memory attention's.
sam21Encoder
image, [3, size, size], to features, high1 and high0.
sam21MasksToImage
masks, [count, size / 4, size / 4], to image, [count, height, width]: the tracker's masks brought to the original frame bilinearly, as the upstream's post_process_masks brings them — straight from a quarter of the square, since SAM 2 resizes a frame without padding it. Reads no weights.
sam21MemoryAttention
features, [dim, grid, grid], memory and its positions, [frames·grid² + pointers, memory.dim], to conditioned, [dim, grid, grid]: one graph for each count of frames and of pointer tokens a tracker holds. The host makes the memory's positions, so this graph reads the temporal table and the pointers' projection for the file to hold them.
sam21MemoryEncoder
features, [dim, grid, grid], and mask, [1, size, size] — logits, or a binary mask when binary, as a frame prompted with points is encoded — to memory, [memory.dim, grid, grid]. The host adds the occlusion embedding where the object is absent, and keeps the memory rounded to bfloat16, as the upstream keeps it; this graph reads the embedding so the file holds it.
sam21Upscale
masks, [count, 4·grid, 4·grid], to high, [count, size, size], bilinearly. Reads no weights.
samGridPositions
[grid · grid, dim]: the position of each cell of the embedding's grid, at its centre, as the upstream's get_dense_pe lays it out and the decoder reads it as tokens.
samInputSize
The upstream's get_preprocess_shape: the longest side to longest, each side rounded.
samMasksToImage
masks, [count, 4·grid, 4·grid], to image, [count, height, width]: up to the encoder's size, cropped to the image as it was prepared, and down to the original — bilinearly, without aligned corners, as the upstream's postprocess_masks does both. Reads no weights.
samPromptTokens
The tokens of request, [samTokenCount(request, layout), dim], into out, for an original image of height by width prepared at size as layout's processor prepares it.
samTokenCount
How many tokens a prompt is, in a layout.
schurSolve
The step for damping, into cameraStep (cameras · cameraSize) and pointStep (points · 3); false where a point's block or the reduced system cannot be factored.
segmentGeometry
Regions from the geometry alone: runs of triangles that share a plane.
selectFrames
The frames a capture keeps, in order, the first among them.
sh1Basis
The three degree-1 basis values for a direction, the band constant folded in, into out.
splatColour
A Gaussian's colour from a direction whose basis is already resolved, clamped at zero.
stabilityScore
The upstream's stability score: the share of the pixels above threshold − offset that are also above threshold + offset — one over the other of the mask thresholded high and thresholded low, since one always lies inside the other. A mask with no pixel above the lower threshold scores NaN, as the upstream's does.
structuralSimilarity
The mean structural similarity of a and b, one value a pixel, over every window.
svd3
The same for the 3 × 3 matrices geometry is full of.
svdN
One-sided Jacobi: m, rows × cols row-major with rows ≥ cols, into u (rows × cols), s (cols, falling) and v (cols × cols), where m = u · diag(s) · vᵀ.
symmetricEigen
Jacobi's eigendecomposition of a symmetric m, n × n: values in falling order and vectors by column, so m · vectors[:, k] = values[k] · vectors[:, k].
triangleResize
src into out by the antialiased bilinear: across first, then down, through an eight-bit intermediate, which is the order both libraries' passes have. precision says whose rounding.
triangulate
The points behind count correspondences, in the first camera's frame, into out (3 each), and how many landed in front of both cameras. A point that did not is left at zero.
visibleGaussians
The Gaussians a camera sees, nearest first, and where each lands.

Types

Frame
A frame this rasteriser writes: single precision where it stands in for a render target, double where the caller is differentiating through it. The fitting loop accumulates its loss in double and a frame quantised to single is a floor under every difference taken through one — measured at about a part in five hundred of the slope, against a part in ten million when the frame is double.
GraphRun

Constants

DEPTH_ANYTHING_2_SMALL
The checkpoint's configuration at its pinned revision.
DEPTH_ANYTHING_3
The two accepted sizes, from each checkpoint's own configuration at its pinned revision.
DEPTH_PIXEL_MEAN
DEPTH_PIXEL_STD
DESCRIPTOR_BYTES
A description is 256 bits, which is 32 bytes.
LOW_PASS
The low-pass the shader adds to the screen variances, in square pixels.
MINIATURE_CASES
One view on a grid its positions are resized to, and two views on the trained grid.
MINIATURE_DEPTH_ANYTHING_2
Depth Anything V2 in miniature: four blocks of 32 channels in four heads, every block tapped, a trained grid of 2, and a neck and head of 16 and 8.
MINIATURE_DEPTH_ANYTHING_3
Four blocks of 32 channels in four heads of 8 — the smallest a two-dimensional rotary embedding divides — with the camera token, query-key norms and rotary embedding from the third block, so one within-view and one across-view block each carry all three; a trained grid of 2, so any other grid is resized; and a head of 16 features.
MINIATURE_MOBILE_SAM
MINIATURE_OWLV2
MINIATURE_OWLV2_IMAGE
The original image's size the boxes are scaled to: wider than tall, padded below.
MINIATURE_OWLV2_MERGES
c+a, ca+t</w> and a+t</w>: tokens 512 to 514, and the start and end of text 515 and 516.
MINIATURE_OWLV2_QUERIES
Three queries as token ids: one ending early, one with the padding token inside it — the "!" a text may hold — and one filling every position.
MINIATURE_SAM_21
MINIATURE_SAM_21_IMAGE
The original frame the prompts are given in, 36 by 48, which SAM 2 resizes to its square — so each axis scales by its own factor — and a point, a box, and both with a second point off the object.
MINIATURE_SAM_21_PROMPTS
MINIATURE_SAM_IMAGE
The original image the prompts are given in: 36 by 48, prepared at 48 by 64 in the encoder's 64 square, so the masks are cropped before they come down to it and coordinates are scaled by 4/3.
MINIATURE_SAM_PROMPTS
A point, which a padding point follows; a box; and two points with a box, which none follows.
MINIATURE_SEED
The seed the tests and the oracle share.
MOBILE_SAM
MOVABLE_COMPONENTS
Something small enough and named: a scene may want to move it.
OWLV2_BASE
OWLV2_PIXEL_MEAN
The pixel mean and deviation an image is normalised by, CLIP's, in 0–1 RGB.
OWLV2_PIXEL_STD
PILLOW_PRECISION
Pillow's fraction bits, and torchvision's eight-bit path's.
PROPOSAL_COMPONENT
The component a capture proposes in. Named for what it is: a proposal, not a thing.
PROPOSAL_SCHEMA
The fields a proposal carries, by stable id.
REACH
How many standard deviations of a splat are drawn, and the largest radius in pixels.
SAM_21_TINY
SAM_PIXEL_MEAN
The pixel mean and deviation the encoder's input is normalised by, in 0–255 RGB.
SAM_PIXEL_STD
SAM_PROMPT
SAM2_PROMPT
SAM2_VIDEO_PROMPT
SAM 2 in a video session: the same encoder, a box as its two corner points first.
SCENERY_COMPONENTS
Drawn and solid and nothing else, which is what an unlabelled region gets.
SH_C1
The l=1 basis constant, as the shader's SH_C1.
SH1_COEFFICIENTS
Nine coefficients a Gaussian, interleaved by basis and then by channel.
SIMILARITY_WINDOW
The side of the block the three statistics are taken over.
TORCHVISION_PRECISION
WALKABLE_COMPONENTS
The same, and a navigation surface on top of it.