DriftCapture · interface
ClipTokenizer
CLIP's tokenizer — byte-level BPE — as Transformers' CLIPTokenizer runs it at the revision the
manifest pins, for OWLv2's text queries.
Explained in DriftCapture.
interface ClipTokenizerimport type { ClipTokenizer } from '@driftengine/capture';In depth
Its whole vocabulary is its merges. A byte's token is the byte's place in GPT-2's order —
the printable bytes first, then the rest — the same byte ending a word is 256 further on, the
merge of rank r makes token 512 + r, and the start and end of text follow the last merge. So a
converted file carries the merges as pairs of token ids and nothing else, and the conversion
checks the upstream's vocab.json says the same.
In the upstream's order: ! and <|endoftext|> are taken out of the raw text as the added
tokens they are — ! is the padding token, 0, wherever it appears; the rest is composed (NFC)
and lowercased a character at a time; <|startoftext|> is then taken out of that; what is left
is split by CLIP's pattern into contractions, runs of letters, single digits and runs of anything
else; each piece's UTF-8 bytes become tokens with the last marked as a word's end, and adjacent
pairs merge lowest rank first, leftmost among equals. A pair listed twice ranks as its last, as
the upstream's map of merges keeps it.
What it leaves out, and why that changes nothing: the upstream also makes each run of whitespace one space, and splits again by GPT-2's pattern after CLIP's. No piece holds whitespace, so the first cannot change a token; and neither pattern crosses whitespace while every piece CLIP's makes is one GPT-2's matches whole, so the second divides nothing. What would make that wrong is a pattern change on either side, which the hand-run parity would show.
Methods
encode
encode(text: string): number[]The text's tokens, between the start and end of text.
| Parameter | Type | Description |
|---|---|---|
text | string |