We’re very excited to present this new method of building powerful, general purpose encoder-decoder models by adapting from pretrained decoder-only LLMs like Gemma 2. To help accelerate further research and allow the community to build on this work, we are excited to release a suite of our T5Gemma checkpoints. The release includes:
Multiple Sizes: Checkpoints for T5-sized models (Small, Base, Large, and XL), the Gemma 2-based models (2B and 9B), as well as an additional model in between T5 Large and T5 XL.
Flexible Configurations: A powerful and efficient unbalanced 9B-2B checkpoint to explore the trade-offs between encoder and decoder size.
Different Training Objectives: Models trained with either PrefixLM or UL2 objectives to provide either state-of-the-art generative performance or representation quality.
We hope these checkpoints will provide a valuable resource for investigating model architecture, efficiency, and performance.
We can't wait to see what you build with T5Gemma. Please see the following links for more information:
Learn about the research behind this project by reading the paper .
Download the models: Find the model weights on Hugging Face and Kaggle .
Explore the models capabilities or fine-tune them for your own use cases with the Colab notebook .
Source link







