Making and shipping
DriftCapture
200 symbols, imported from @driftengine/capture.
Explained in DriftCapture.
Classes
Interfaces
BlockNormals- The normal equations in blocks: cameras, points of three, and one link per observation.
BrowserClip- A browser's clip, and how to let it go.
CaptureFrameCaptureSceneClipTokenizer- CLIP's tokenizer — byte-level BPE — as Transformers'
CLIPTokenizerruns it at the revision the manifest pins, for OWLv2's text queries. ClipTowerConfigCollisionOptionsCollisionSource- A mesh as
meshShapetakes it, and what had to go for it to be one. DecimateOptionsDelightOptionsDelightOut- What comes back, one entry per vertex of the mesh handed in.
DelightView- One view of the surface: a linear frame and where it was taken from.
DepthAnything2ConfigDepthAnything3ConfigDepthEstimateDepthEstimatorDepthViewDescriptors- The descriptions of a frame's features:
BYTESbytes each, bit by comparison. EntityProposal- What a scene could make of a region.
FeatureSet- The features of one frame, in the arrays a caller owns.
FitOptionsFrameChoiceFrameSize- A clip's frames, and which of them a capture keeps.
FrameSourceFuseOptionsGaussianGradients- Where the gradients go: one entry per parameter of every Gaussian.
GaussianSet- Where a Gaussian lands on screen: the projection
@driftengine/splats' shader performs, in JavaScript, and the order a frame draws them in. HeadConfigHieraConfigLeastSquares- A least-squares problem: how many residuals, how many parameters, and both at a point.
LightLitSceneMarchOptionsMaskView- Where a model is prompted, and what came back for each prompt.
MeshHitMobileSamConfigOptimiseOptionsOwlv2ConfigOwlv2DetectionOwlv2PatchesPoseOptionsPoseResultPreparedFrameProjected- One Gaussian as the screen sees it.
PropHulls- A prop's hulls, and how much empty space they added.
ProposeOptionsRasterCameraRawFrame- A frame as a host hands it over: RGBA bytes, and the size they are.
Region- What a region is, once something has decided where it is.
RelativePoseOutRelativePoseResultSam21ConfigSam21Frame- A frame's result: every object's mask logits at a quarter of the encoder's size, and its score.
Sam21MemoryConfigSam21VideoSam21Weights- The weights each graph reads: a converted file's graphs, or a checkpoint for all four.
SamDecoderConfigSamPromptSamPromptLayout- Where a checkpoint keeps its prompt encoder, and how its processor places a prompt: SAM's and MobileSAM's, or Transformers' SAM 2.
SegmentOptionsSurfaceView- One view's opinion: its depth, how well covered each pixel was, and where it stood.
TestBoxTestCamera- Analytic scenes rendered on the CPU, so a capture's stages can be held to geometry that is known before anything runs.
TestPlaneTestSceneTinyVitConfigVitConfigVolume- A field of samples over a box of the world,
xfastest, theny, thenz.
Functions
accumulateGradients- The gradients of a render against
dPixels— the loss's derivative by each rendered channel, four a pixel — added intoout. The render is repeated here rather than taken as an argument, because the backward walk needs what each Gaussian contributed on the way. areaResizesrc,width × heightpixels ofchannelsbytes, intooutatoutWidth × outHeight. Only shrinking is defined, which is what a frame's preparation asks for.browserFrameSource- The clip's frames, addressed at the clip's rate.
frameRateis the fallback where the browser cannot say what the rate is. captureFile- Every stage of
sceneas one.drft. cholesky- The lower triangle
lwithl · lᵀ = m, or false wheremis not positive definite — which is how a caller learns that its normal equations are singular, rather than by finding a NaN later. choleskySolve- Solves
l · lᵀ · x = bin place overx, given the Cholesky factor. clipTokenizer- A tokenizer over
merges,[count, 2]token ids in rank order. collisionMeshmesh's triangles, less the ones with no area, renumbered onto the vertices that remain.createDelightOutcreateDepthEstimator- An estimator over one model's weights, running its graphs through
run. createFeatureSetcreateGradientscreateVolumedecimatemeshsimplified until it holds at mostbudgettriangles.decodeCamera- One view's camera from its pose encoding:
worldToCamera, 3 × 4 row-major — the inverse of the camera-to-world rotation and translation the decoder predicts — andintrinsics, 3 × 3 row-major, with focal lengths from the two fields of view and the principal point at the centre. decodeDepth- Depth and confidence from one view's logits,
[2, height, width]:expof the first channel, and one more thanexpof the second, into the caller's maps. delightmesh's material, as seen inviews.depthAnything2- The graph for one image of
height × widthpixels, each a multiple of the patch, normalised by ImageNet's mean and deviation: inputimage,[3, height, width]; outputdepth,[1, height, width]. depthAnything3- The graph for
viewsimages ofheight × widthpixels, each a multiple of 14, normalised by ImageNet's mean and deviation. Inputs areimage0… as[3, height, width]; outputs arelogits0… as[2, height, width]andposeas[views, 9]. The first view is the reference. describeFeatures- A description of every feature in
features, 32 bytes each. detectFeatures- The strongest corners of
image, RGBA atwidth × height, intoout— at mostbudgetof them, strongest first. eachFrame- Every frame in order, each read over the last into one buffer the visitor must not keep.
estimatePoses- The clip's camera path, into
out: 3 × 4 world-to-camera a frame. fuseDepth- Every view's surface written into
volume, which accumulates rather than replaces. imageLoss- The squared difference between a render and a frame, and the render's own gradient.
interpolate- A per-vertex attribute at a hit, by its barycentric weights.
labelMasks- A name for each mask index, from a detector's boxes.
levenbergMarquardt- Levenberg–Marquardt over
x, in place: the Gauss–Newton step withdamping · diag(JᵀJ)added, the damping raised where a step would cost more and lowered where it pays, which is the whole reason it is not Gauss–Newton — a step that overshoots is rejected rather than taken, and a start far from the answer is where that decides between converging and diverging. liftMasks- Masks from several views lifted onto the mesh, by what most of the views that saw a triangle say.
longestSideSize- The size a frame is resized to, its longest side at
size, rounded up at a half. lookAt- A world-to-camera for an eye looking at a target, with the world's up.
marchVolumevolume's surface, welded so that neighbours share vertices exactly.matchFeatures- The matches between two frames' descriptions into
out, two indices each, and how many there are. A match is mutual and clears Lowe'sratio, which defaults to 0.8. miniatureCheckpoint- Every tensor the upstream's backbone, head and camera decoder hold, named as the checkpoints are.
miniatureCheckpoint2- Every tensor Transformers' Depth Anything holds, in the older names its checkpoints keep.
miniatureImages- Seeded images for
viewsviews, normalised as the model expects,[3, height, width]each. miniatureOwlv2Checkpoint- Every tensor the upstream's
Owlv2ForObjectDetectionholds, named as its checkpoints are. miniatureSam21Checkpoint- Every tensor the upstream's
Sam2VideoModelholds, named as its checkpoints are. miniatureSamCheckpoint- Every tensor the upstream's encoder, prompt encoder and mask decoder hold.
mobileSamDecoderembeddingandprompt,[tokens, dim]— absent whentokensis zero — andmaskwhenrefineis set, tomasks,[4, 4·grid, 4·grid], andquality,[1, 4].mobileSamEncoderimage,[3, size, size], toembedding,[neck, grid, grid].nearestHit- The nearest triangle a ray meets, by Möller–Trumbore over every one of them.
optimiseGaussians- A clip fitted to Gaussians. Nothing here is the caller's array; everything is freshly cut.
owlv2BoxBias[x, y, w, h]for every cell, row by row: the logits of its far corner and of the cell's size.owlv2Detections- The detections above
thresholdamonglogits,[cells, count], withboxes,[cells, 4]as the image graph answers them, for an original image ofheightbywidth. owlv2Imageimage,[3, size, size], to five values a patch:classes,[cells, text.dim], the class embedding before it is normalised;shiftandscale,[cells, 1], the logit's shift and its scale before the ELU;boxes,[cells, 4], centre and size in the image's fractions; andobjectness,[cells, 1], a logit.owlv2Logitsout,[cells, count]: every patch's logit for every one ofcountqueries,[count, width]as the text graph answers them.owlv2Texttokens,[queries · positions], andends,[queries], asowlv2Tokensmakes them, toqueries,[queries, text.dim]: each query's embedding before it is normalised.owlv2Tokens- OWLv2's text queries as its text graph takes them:
tokens,[queries · positions], each query padded with zeros topositions, andends,[queries], the row of each query's end of text in that list — the upstream pools the row of the largest token, which is that one. A query longer thanpositionsis refused, as the upstream's position table refuses it. prepareDepthFramergba,width × heightpixels of four bytes, as Depth Anything 3 takes it atsize, in whole patches ofpatch.prepareOwlv2Framergbaas OWLv2's processor prepares it:[3, square, square]at the model's size.prepareSam2Framergbaas SAM 2's processor prepares it:[3, square, square], each axis scaled on its own, and normalised bymeananddeviationin 0–1.prepareSamFramergbaas MobileSAM's predictor prepares it:[3, square, square], the frame itself in the top left at the size this answers, and zeros after it.projectGaussian- One Gaussian projected, or null where it is behind the camera or too flat to draw.
promptGrid- A lattice of prompt points over an image, seeded.
propHullsmeshdecomposed into convex hulls the container can hold.proposalScene- The proposals as a scene, keyed by field id — the form
serializeWorldwrites. proposeEntities- Each region as a proposal.
rasteriseGaussianssetthroughcameraintoout, four floats a pixel: linear colour, premultiplied as it composites, and the alpha the cloud covered the pixel with.readProposals- The proposals a scene carries, in the order they were written.
relativePose- The pose of the second view relative to the first, from
countcorrespondences in pixels.randomis the caller's seeded generator: a capture answers the same twice. renderDepth- The cloud's depth and coverage, one of each per pixel.
renderLitsceneseen fromcamera, intooutas linear premultiplied RGBA.renderTestScenescenethroughcameraintoout, RGBA, and the distance along the camera's own axis intodepthwhere one is given — zero where the ray met nothing.sam21Decoderembedding,high1,high0andprompt,[tokens, dim]— absent whentokensis zero — andmaskwhenrefineis set, tomasks,quality,objectandpointers. WithnoMemorythe decoder addsno_memory_embeddingto the features itself, as an image with nothing remembered is decoded; otherwise the features are the memory attention's.sam21Encoderimage,[3, size, size], tofeatures,high1andhigh0.sam21MasksToImagemasks,[count, size / 4, size / 4], toimage,[count, height, width]: the tracker's masks brought to the original frame bilinearly, as the upstream'spost_process_masksbrings them — straight from a quarter of the square, since SAM 2 resizes a frame without padding it. Reads no weights.sam21MemoryAttentionfeatures,[dim, grid, grid],memoryand itspositions,[frames·grid² + pointers, memory.dim], toconditioned,[dim, grid, grid]: one graph for each count of frames and of pointer tokens a tracker holds. The host makes the memory's positions, so this graph reads the temporal table and the pointers' projection for the file to hold them.sam21MemoryEncoderfeatures,[dim, grid, grid], andmask,[1, size, size]— logits, or a binary mask whenbinary, as a frame prompted with points is encoded — tomemory,[memory.dim, grid, grid]. The host adds the occlusion embedding where the object is absent, and keeps the memory rounded to bfloat16, as the upstream keeps it; this graph reads the embedding so the file holds it.sam21Upscalemasks,[count, 4·grid, 4·grid], tohigh,[count, size, size], bilinearly. Reads no weights.samGridPositions[grid · grid, dim]: the position of each cell of the embedding's grid, at its centre, as the upstream'sget_dense_pelays it out and the decoder reads it as tokens.samInputSize- The upstream's
get_preprocess_shape: the longest side tolongest, each side rounded. samMasksToImagemasks,[count, 4·grid, 4·grid], toimage,[count, height, width]: up to the encoder's size, cropped to the image as it was prepared, and down to the original — bilinearly, without aligned corners, as the upstream'spostprocess_masksdoes both. Reads no weights.samPromptTokens- The tokens of
request,[samTokenCount(request, layout), dim], intoout, for an original image ofheightbywidthprepared atsizeaslayout's processor prepares it. samTokenCount- How many tokens a prompt is, in a layout.
schurSolve- The step for
damping, intocameraStep(cameras · cameraSize) andpointStep(points · 3); false where a point's block or the reduced system cannot be factored. segmentGeometry- Regions from the geometry alone: runs of triangles that share a plane.
selectFrames- The frames a capture keeps, in order, the first among them.
sh1Basis- The three degree-1 basis values for a direction, the band constant folded in, into
out. splatColour- A Gaussian's colour from a direction whose basis is already resolved, clamped at zero.
stabilityScore- The upstream's stability score: the share of the pixels above
threshold − offsetthat are also abovethreshold + offset— one over the other of the mask thresholded high and thresholded low, since one always lies inside the other. A mask with no pixel above the lower threshold scores NaN, as the upstream's does. structuralSimilarity- The mean structural similarity of
aandb, one value a pixel, over every window. svd3- The same for the 3 × 3 matrices geometry is full of.
svdN- One-sided Jacobi:
m,rows × colsrow-major withrows ≥ cols, intou(rows × cols),s(cols, falling) andv(cols × cols), wherem = u · diag(s) · vᵀ. symmetricEigen- Jacobi's eigendecomposition of a symmetric
m,n × n:valuesin falling order andvectorsby column, som · vectors[:, k] = values[k] · vectors[:, k]. triangleResizesrcintooutby the antialiased bilinear: across first, then down, through an eight-bit intermediate, which is the order both libraries' passes have.precisionsays whose rounding.triangulate- The points behind
countcorrespondences, in the first camera's frame, intoout(3 each), and how many landed in front of both cameras. A point that did not is left at zero. visibleGaussians- The Gaussians a camera sees, nearest first, and where each lands.
Types
Frame- A frame this rasteriser writes: single precision where it stands in for a render target, double where the caller is differentiating through it. The fitting loop accumulates its loss in double and a frame quantised to single is a floor under every difference taken through one — measured at about a part in five hundred of the slope, against a part in ten million when the frame is double.
GraphRun
Constants
DEPTH_ANYTHING_2_SMALL- The checkpoint's configuration at its pinned revision.
DEPTH_ANYTHING_3- The two accepted sizes, from each checkpoint's own configuration at its pinned revision.
DEPTH_PIXEL_MEANDEPTH_PIXEL_STDDESCRIPTOR_BYTES- A description is 256 bits, which is 32 bytes.
LOW_PASS- The low-pass the shader adds to the screen variances, in square pixels.
MINIATURE_CASES- One view on a grid its positions are resized to, and two views on the trained grid.
MINIATURE_DEPTH_ANYTHING_2- Depth Anything V2 in miniature: four blocks of 32 channels in four heads, every block tapped, a trained grid of 2, and a neck and head of 16 and 8.
MINIATURE_DEPTH_ANYTHING_3- Four blocks of 32 channels in four heads of 8 — the smallest a two-dimensional rotary embedding divides — with the camera token, query-key norms and rotary embedding from the third block, so one within-view and one across-view block each carry all three; a trained grid of 2, so any other grid is resized; and a head of 16 features.
MINIATURE_MOBILE_SAMMINIATURE_OWLV2MINIATURE_OWLV2_IMAGE- The original image's size the boxes are scaled to: wider than tall, padded below.
MINIATURE_OWLV2_MERGES- c+a, ca+t</w> and a+t</w>: tokens 512 to 514, and the start and end of text 515 and 516.
MINIATURE_OWLV2_QUERIES- Three queries as token ids: one ending early, one with the padding token inside it — the "!" a text may hold — and one filling every position.
MINIATURE_SAM_21MINIATURE_SAM_21_IMAGE- The original frame the prompts are given in, 36 by 48, which SAM 2 resizes to its square — so each axis scales by its own factor — and a point, a box, and both with a second point off the object.
MINIATURE_SAM_21_PROMPTSMINIATURE_SAM_IMAGE- The original image the prompts are given in: 36 by 48, prepared at 48 by 64 in the encoder's 64 square, so the masks are cropped before they come down to it and coordinates are scaled by 4/3.
MINIATURE_SAM_PROMPTS- A point, which a padding point follows; a box; and two points with a box, which none follows.
MINIATURE_SEED- The seed the tests and the oracle share.
MOBILE_SAMMOVABLE_COMPONENTS- Something small enough and named: a scene may want to move it.
OWLV2_BASEOWLV2_PIXEL_MEAN- The pixel mean and deviation an image is normalised by, CLIP's, in 0–1 RGB.
OWLV2_PIXEL_STDPILLOW_PRECISION- Pillow's fraction bits, and torchvision's eight-bit path's.
PROPOSAL_COMPONENT- The component a capture proposes in. Named for what it is: a proposal, not a thing.
PROPOSAL_SCHEMA- The fields a proposal carries, by stable id.
REACH- How many standard deviations of a splat are drawn, and the largest radius in pixels.
SAM_21_TINYSAM_PIXEL_MEAN- The pixel mean and deviation the encoder's input is normalised by, in 0–255 RGB.
SAM_PIXEL_STDSAM_PROMPTSAM2_PROMPTSAM2_VIDEO_PROMPT- SAM 2 in a video session: the same encoder, a box as its two corner points first.
SCENERY_COMPONENTS- Drawn and solid and nothing else, which is what an unlabelled region gets.
SH_C1- The l=1 basis constant, as the shader's
SH_C1. SH1_COEFFICIENTS- Nine coefficients a Gaussian, interleaved by basis and then by channel.
SIMILARITY_WINDOW- The side of the block the three statistics are taken over.
TORCHVISION_PRECISIONWALKABLE_COMPONENTS- The same, and a navigation surface on top of it.