Introducing TexTok: Text-Conditioned Tokenization
TexTok architecture: text embeddings guide both tokenization and detokenization.
TexTok architecture: text embeddings guide both tokenization and detokenization.
  • Injects text embeddings into the tokenization process to provide semantic context.
  • Text guides the tokenizer to focus on fine-grained visual details instead of high-level concepts.
  • Simplifies semantic learning, freeing token space for high-fidelity reconstruction.
Semantic Scaffolding

By offloading the burden of semantic understanding to text, TexTok dedicates more model capacity to capturing intricate visual details, enabling higher compression without sacrificing quality.