uluka.labs AWS portal

Our models

Each model tries our method on a different design. Uluka 26B and 7B are proofs of concept, not finished products. Uluka 320B, planned next, should save the most.

Proof of concept

Uluka 26B

Compressed from Google's Gemma 4 26B

51.6 GB to 8.7 GB, 5.9× smaller

  • Averages 69.9 on four benchmarks against 70.4 for Google's 4-bit version, using 44% less memory
  • Reads text and images
  • Runs on a single NVIDIA GPU with 12 GB or more

See the results

Proof of concept

Uluka 7B

Compressed from AI2's OLMoE-1B-7B

13.8 GB to 2.8 GB, 4.9× smaller

  • Keeps about 98% of the original's accuracy across seven benchmarks, corrected for chance
  • Runs on a laptop with no dedicated GPU: about 6 tokens a second on an Intel Core Ultra 5
  • We built it to check the method also works on a second model design. Speed was not a goal.

Planned

Uluka 320B

To be compressed from Z.ai's GLM-5.3-Flash

Target: about 85 GB on a single 96 GB GPU

  • GLM-5.3-Flash is about 12 times the size of Gemma 4 26B
  • Its 4-bit formats need two to four 96 GB GPUs. We aim to fit it on one. We project that would cut GPU cost by 59% to 80%.
  • Not built yet. These are targets, not results.

How larger models save more