Why Uluka
A model has to fit in a GPU's memory, and GPUs with more memory cost much more. We compress the model so it fits on a cheaper GPU.
GPU memory sets the price
A model runs inside a GPU's memory, so its size decides which GPU you rent. The price climbs quickly with memory: a 16 GB GPU costs $0.53 an hour and an 80 GB GPU costs $6.88.
Google's 4-bit Gemma 4 26B is 15.6 GB. That leaves under 0.5 GB free on a 16 GB GPU, so it needs the 24 GB one. Uluka 26B is 8.7 GB and fits on the 16 GB GPU.
Price per hour, by GPU memory
16 GBT4. Fits Uluka 26B, 8.7 GB
24 GBL4. Fits Google's 4-bit, 15.6 GB
48 GBL40S. Room for larger models
80 GBH100. Fits the original, 51.6 GB
Large models need several GPUs
A model too big for one GPU has to be split across several, and the bill multiplies. GLM-5.3-Flash from Z.ai is a 320B model, about 12 times the size of Gemma 4 26B. Its usual 4-bit formats (MXFP4, NVFP4, Q4) are 170 to 204 GB and need two to four 96 GB GPUs.
Our custom compression architecture should get the model to about 85 GB, small enough for one GPU.
GLM-5.3-Flash in each format
OriginalBF16, Z.ai
FP8Z.ai
NVFP4NVIDIA
Q4UD-Q4_K_XL, Unsloth
MXFP4calculated
Q3UD-IQ3_XXS, Unsloth
Uluka 320Bour target
Where we are
What works today, and what we are building next.
- WorkingCompression and runtime work on two proof-of-concept models: Uluka 26B and Uluka 7B
- In progressServing more users at once, and using less memory in long conversations
- PlannedUluka 320B, compressed from GLM-5.3-Flash, and support for hardware beyond NVIDIA GPUs