Rendering
Custom passes and compute
Add a draw pass of your own, fill targets it owns before the frame opens, dispatch compute on WebGPU, and keep a pass right on both backends.
npm install @driftengine/core
The renderer's verbs cover what most games draw, and a pass of your own covers the rest: a shader written for one effect, a package's own renderer, a compute kernel that sorts, bins or simulates on the device. The engine's optional packages draw through the same door, so the Gaussian splats and the GPU-driven pipeline are passes registered exactly as yours would be.
A pass
/** A full-screen wash of colour over the frame, on WebGPU, as the smallest pass that draws. */
const WASH_WGSL = /* wgsl */ `
@vertex fn vs(@builtin(vertex_index) i: u32) -> @builtin(position) vec4f {
let corner = vec2f(f32((i << 1u) & 2u), f32(i & 2u));
return vec4f(corner * 2.0 - 1.0, 0.0, 1.0);
}
@fragment fn fs() -> @location(0) vec4f { return vec4f(0.1, 0.0, 0.2, 0.15); }
`;
export function washPass(): PassDefinition {
let pipeline: GPURenderPipeline | null = null;
return {
label: 'wash',
init(device) {
if (device.backend !== 'webgpu') return;
const module = device.device.createShaderModule({ code: WASH_WGSL });
pipeline = device.device.createRenderPipeline({
layout: 'auto',
vertex: { module, entryPoint: 'vs' },
fragment: {
module,
entryPoint: 'fs',
targets: [
{
format: device.format,
blend: {
color: { srcFactor: 'src-alpha', dstFactor: 'one-minus-src-alpha' },
alpha: SCENE_ALPHA_COVERS,
},
},
],
},
depthStencil: {
format: device.depthFormat,
depthWriteEnabled: false,
depthCompare: 'always',
},
multisample: { count: device.samples },
});
},
draw(ctx) {
if (ctx.backend !== 'webgpu' || pipeline === null) return;
ctx.pass.setPipeline(pipeline);
ctx.pass.draw(3);
},
};
}A PassDefinition has a label and up to four callbacks:
init(device)builds pipelines, buffers and textures, once, when the pass is registered.prepare(ctx)fills a target the pass owns, each frame, before the frame's own target opens: a shadow map of its own, a blur pyramid, a picking buffer.draw(ctx)draws into the frame, at the point the frame callsdrawPass.dispose(device)releases whatinitbuilt.
reads declares the frame attachments the pass samples, such as depth, so the frame keeps them for
it. What a pass writes is wherever the frame is drawing; a pass cannot declare other writes into the
frame, which is why prepare exists for targets it owns.
The device and both contexts are tagged by backend, because a pass that draws on both has a
different shader and a different buffer layout on each and branches either way. On WebGPU the
device carries the frame's colour format, depthFormat and samples, which a pipeline cannot be
built without, and whether the frame is reconstructed. On WebGL2 it carries the context.
/** Registered once; drawn wherever the frame calls `drawPass`; let go when done. */
export function useWash(renderer: RendererApi): { draw(): void; release(): void } {
const handle = renderer.registerPass(washPass());
return {
draw: () => renderer.drawPass(handle),
release: () => renderer.unregisterPass(handle),
};
}registerPass(definition) returns a handle, drawPass(handle) draws it inside the frame, and
unregisterPass(handle) calls dispose. A handle kept after it is unregistered draws nothing; it
never draws whatever took its slot.
The rules a pass keeps
- It allocates nothing in
prepareordraw. They run every frame, and the engine's own rule about per-frame code binds them. - It leaves the context as it found it. On WebGL2: the program, the vertex array, blend and depth
state and the viewport;
preparealso leaves the default framebuffer bound. On WebGPU only the viewport and scissor carry over between draws. - It corrects its projection. A pass takes its camera from you, so it never sees the corrected
matrix the renderer builds. Pre-multiply by
device.clipCorrectionif your WGSL was generated from GLSL with the engine's generator, and bydevice.depthCorrectionif you wrote it by hand. The wrong one mirrors the world vertically, and mirrors the triangle winding with it, so it looks like a culling problem. - It grades only when the frame does not.
ctx.outputTransformis 0 when the end of the frame applies the tone curve, and otherwise says which curve this pass must apply itself, withctx.outputExposure.OUTPUT_TRANSFORM_GLSLis the engine's own curve, so a pass applies the same one. - It jitters with a reconstructed frame. With DriftTR on, a pass drawing the world applies the
frame's jitter,
ctx.jitter, withjitterClip, and writes its depth into the frame's attachment; see DriftTR. - It says what its blending does to the scene's alpha. The scene's alpha holds how much of each
pixel is still the opaque surface, and the composite darkens only that share with ambient
occlusion, so a translucent draw is not darkened by a corner it hides. A pass drawing something
blended over the world gives its alpha
SCENE_ALPHA_COVERS(it hides what is behind by its alpha),SCENE_ALPHA_TRANSMITS(its alpha is what gets through) orSCENE_ALPHA_KEEPS(it adds light to the surface); on WebGL2,blendCovering,blendTransmittingorblendKeepingset the same thing beside the colour factors. An opaque draw writes 1. - It sizes its targets from the scene.
renderer.sceneWidthandsceneHeightare what the frame is drawn at, which is smaller than the canvas under reconstruction.
On WebGPU, prepare is also handed the frame's environment probe and, with DriftRay on, the world's
composed distance field, one frame behind, for a pass that wants to light or trace its own work the
way the frame does.
Reading the frame
A pass cannot sample the frame it is drawing into, so it reads copies, and declares which:
reads: ['depthSnapshot']hands it the frame's depth as it stands, one copy a frame, assceneDepth: on WebGPU in thePrepareContext, a single-sampler32floattexture to read withtextureLoador a non-filtering sampler, and on WebGL2 in thePassContext. It holds the depth buffer's own values, reversed as the buffer is: 1 at the near plane, 0 at the far one.reads: ['colorSnapshot']hands it the frame's colour with everything drawn before it, copied afresh at each such pass's draw, assceneColor: on WebGPU in thePrepareContext, at the scene's size and format, and on WebGL2 in thePassContext. A distortion samples the scene behind it with it; a full-screen pass drawn last reads the whole frame and writes over it, in scene light, before the tone curve: a radial blur, a flash, a filter over the finished picture.
The views on WebGPU arrive in prepare, since a bind group is built before the frame's pass opens,
and are filled at the draw. Both copies are null without screenEffects, which is what keeps a
scene target to copy from.
Under DriftTR a pass that reads the colour draws after the upscale, over the reconstructed
picture and after every blended draw, whatever order it was asked in: otherwise it would read a frame
without the blended draws and then be covered by them. It draws at the output's size and the
reconstruction's format, which ctx.format names there, so it keeps a pipeline for that target. Its
sceneDepth stays the render's opaque depth: coarser by the ratio above a ratio of 1, and a blended
draw is in no depth either way.
Compute
/** A buffer of counters cleared on the device every frame, as the smallest compute definition. */
const CLEAR_WGSL = /* wgsl */ `
@group(0) @binding(0) var<storage, read_write> counts: array<u32>;
@compute @workgroup_size(64) fn main(@builtin(global_invocation_id) id: vec3u) {
if (id.x < arrayLength(&counts)) { counts[id.x] = 0u; }
}
`;
/** `GPUBufferUsage.STORAGE`, spelled as its value: the global exists only in a browser. */
const STORAGE = 0x0080;
export function clearCounters(size: number): ComputeDefinition {
let pipeline: GPUComputePipeline | null = null;
let group: GPUBindGroup | null = null;
return {
label: 'clear counters',
init({ device }) {
const buffer = device.createBuffer({ size: size * 4, usage: STORAGE });
pipeline = device.createComputePipeline({
layout: 'auto',
compute: { module: device.createShaderModule({ code: CLEAR_WGSL }), entryPoint: 'main' },
});
group = device.createBindGroup({
layout: pipeline.getBindGroupLayout(0),
entries: [{ binding: 0, resource: { buffer } }],
});
},
dispatch({ pass }) {
if (pipeline === null || group === null) return;
pass.setPipeline(pipeline);
pass.setBindGroup(0, group);
pass.dispatchWorkgroups(Math.ceil(size / 64));
},
};
}
/** Only where the backend has compute at all; WebGL2 does not. */
export function startCounters(renderer: RendererApi): (() => void) | null {
if (!renderer.computeSupported) return null;
const handle = renderer.registerCompute(clearCounters(4096));
return () => renderer.dispatchCompute(handle);
}A ComputeDefinition is the same idea for the stage before drawing: init builds pipelines and
buffers from the device, and dispatch(ctx) sets a pipeline and bind groups and calls
dispatchWorkgroups on the compute pass the engine opened for it. The workgroup count is yours,
since it comes from data the definition owns. registerCompute(definition) returns a handle,
dispatchCompute(handle) runs it, and unregisterCompute releases it.
Compute is WebGPU only: WebGL2 has no compute stage. renderer.computeSupported says which you
have, and registerCompute refuses in words where it is false. A dispatch is submitted at once, so
its result is visible to everything recorded after it, including a draw pass in the same frame.
Testing a compute shader
WGSL arithmetic can be wrong in ways no picture shows. The engine tests its own kernels by running
them on a real device and reading the buffers back against a TypeScript reference, with
openGpuCompute from @driftengine/core/scripts/gpuCompute.mjs. The same harness works for a game's own kernels;
it serves its page from a loopback server, since navigator.gpu needs a secure context.