Proof of concept
Uluka 26B
Compressed from Google's Gemma 4 26B
51.6 GB to 8.7 GB, 5.9× smaller
- Averages 69.9 on four benchmarks against 70.4 for Google's 4-bit version, using 44% less memory
- Reads text and images
- Runs on a single NVIDIA GPU with 12 GB or more
See the results
Proof of concept
Uluka 7B
Compressed from AI2's OLMoE-1B-7B
13.8 GB to 2.8 GB, 4.9× smaller
- Keeps about 98% of the original's accuracy across seven benchmarks, corrected for chance
- Runs on a laptop with no dedicated GPU: about 6 tokens a second on an Intel Core Ultra 5
- We built it to check the method also works on a second model design. Speed was not a goal.
Planned
Uluka 320B
To be compressed from Z.ai's GLM-5.3-Flash
Target: about 85 GB on a single 96 GB GPU
- GLM-5.3-Flash is about 12 times the size of Gemma 4 26B
- Its 4-bit formats need two to four 96 GB GPUs. We aim to fit it on one. We project that would cut GPU cost by 59% to 80%.
- Not built yet. These are targets, not results.
How larger models save more