DriftCapture · interface

ClipTokenizer

CLIP's tokenizer — byte-level BPE — as Transformers' CLIPTokenizer runs it at the revision the manifest pins, for OWLv2's text queries.

Explained in DriftCapture.

interface ClipTokenizer
import type { ClipTokenizer } from '@driftengine/capture';

In depth

Its whole vocabulary is its merges. A byte's token is the byte's place in GPT-2's order — the printable bytes first, then the rest — the same byte ending a word is 256 further on, the merge of rank r makes token 512 + r, and the start and end of text follow the last merge. So a converted file carries the merges as pairs of token ids and nothing else, and the conversion checks the upstream's vocab.json says the same.

In the upstream's order: ! and <|endoftext|> are taken out of the raw text as the added tokens they are — ! is the padding token, 0, wherever it appears; the rest is composed (NFC) and lowercased a character at a time; <|startoftext|> is then taken out of that; what is left is split by CLIP's pattern into contractions, runs of letters, single digits and runs of anything else; each piece's UTF-8 bytes become tokens with the last marked as a word's end, and adjacent pairs merge lowest rank first, leftmost among equals. A pair listed twice ranks as its last, as the upstream's map of merges keeps it.

What it leaves out, and why that changes nothing: the upstream also makes each run of whitespace one space, and splits again by GPT-2's pattern after CLIP's. No piece holds whitespace, so the first cannot change a token; and neither pattern crosses whitespace while every piece CLIP's makes is one GPT-2's matches whole, so the second divides nothing. What would make that wrong is a pattern change on either side, which the hand-run parity would show.

Methods

encode

encode(text: string): number[]

The text's tokens, between the start and end of text.

ParameterTypeDescription
textstring