-
Injects text embeddings into the tokenization process to provide semantic context.
-
Text guides the tokenizer to focus on fine-grained visual details instead of high-level concepts.
-
Simplifies semantic learning, freeing token space for high-fidelity reconstruction.
Semantic Scaffolding
By offloading the burden of semantic understanding to text, TexTok dedicates more model capacity to capturing intricate visual details, enabling higher compression without sacrificing quality.