What a smaller model saves
A smaller model fits on a cheaper GPU. In our first test that cut the monthly bill by 35%, and we project more for much larger models.
The same model on a cheaper GPU
Google's 4-bit Gemma 4 26B is too big for the 16 GB GPU, so it needs the 24 GB one. Uluka 26B fits the 16 GB GPU.
Each server runs one copy of the model. The costs and the saving below scale with the number of servers.
Google's 4-bit version
15.6 GB
- GPU it needs
- 24 GB (L4)
- GPU price
- $0.80 an hour
$588 a month
$0.80 an hour × 730 hours
Uluka 26B
8.7 GB
- GPU it needs
- 16 GB (T4)
- GPU price
- $0.53 an hour
$384 a month
$0.53 an hour × 730 hours
Uluka saves
$20435% cheaper
a month, or $2,442 a year, against Google's 4-bit setup
GPU prices are Amazon Web Services on-demand, US East, September 2026, for 730 hours a month. This is hourly price only. The two GPUs run at different speeds, so cost per request is not the same as cost per hour.
Why the setups differ
A model has to fit inside its GPU, so each version of Gemma 4 26B needs a different one.
| GPU | Original 51.6 GB | Google 4-bit 15.6 GB | Uluka 26B 8.7 GB |
|---|---|---|---|
| 12 GBRTX 3060, RTX 4070 SUPER | ×Does not fit | ×Does not fit | ✓Fits |
| 16 GBT4 | ×Does not fit | ×Does not fit | ✓Fits |
| 24 GBL4, RTX 4090 | ×Does not fit | ✓Fits | ✓Fits |
| 32 GBRTX 5090 | ×Does not fit | ✓Fits | ✓Fits |
| 48 GBL40S | ×Does not fit | ✓Fits | ✓Fits |
| 80 GBA100, H100 | ✓Fits | ✓Fits | ✓Fits |
A model counts as fitting if 1.5 GB is left free for it to work.
A much larger model
At 26B the saving is modest: community builds of 11 to 14 GB already fit on the 16 GB GPU. The saving should grow with model size. GLM-5.3-Flash is a 320B model, and its usual 4-bit formats need two to four 96 GB GPUs. A version that fits on one should cut GPU cost by more than half. That is the target for Uluka 320B.
Projected cost per server, per month
| GLM-5.3-Flash version | GPUs needed (96 GB each) | Cost per month* | Uluka 320B would save per month* |
|---|---|---|---|
| Uluka 320Bour target, about 85 GB | 1 | $2,455 projected* | – |
| MXFP4calculated, 170 GB | 2 | $6,049 | $3,594 (59%) |
| Q3UD-IQ3_XXS, Unsloth, 120 GB | 2 | $6,049 | $3,594 (59%) |
| NVFP4NVIDIA, 204 GB | 4 | $12,098 | $9,643 (80%) |
| Q4UD-Q4_K_XL, Unsloth, 200 GB | 4 | $12,098 | $9,643 (80%) |
*Projected: Uluka 320B is not built yet. Prices are Amazon Web Services g7e instances, on-demand, US East, September 2026: $3.36 an hour for one GPU, $8.29 for two, $16.57 for four. Versions that need the same number of GPUs cost the same. MXFP4 is calculated from its size, because we found no published build. Speed is not compared.