uluka.labs AWS portal

What a smaller model saves

A smaller model fits on a cheaper GPU. In our first test that cut the monthly bill by 35%, and we project more for much larger models.

The same model on a cheaper GPU

Google's 4-bit Gemma 4 26B is too big for the 16 GB GPU, so it needs the 24 GB one. Uluka 26B fits the 16 GB GPU.

Each server runs one copy of the model. The costs and the saving below scale with the number of servers.

Google's 4-bit version

15.6 GB

GPU it needs
24 GB (L4)
GPU price
$0.80 an hour

$588 a month

$0.80 an hour × 730 hours

Uluka 26B

8.7 GB

GPU it needs
16 GB (T4)
GPU price
$0.53 an hour

$384 a month

$0.53 an hour × 730 hours

Uluka saves

$20435% cheaper

a month, or $2,442 a year, against Google's 4-bit setup

GPU prices are Amazon Web Services on-demand, US East, September 2026, for 730 hours a month. This is hourly price only. The two GPUs run at different speeds, so cost per request is not the same as cost per hour.

Why the setups differ

A model has to fit inside its GPU, so each version of Gemma 4 26B needs a different one.

Which forms of Gemma 4 26B fit on each GPU
GPUOriginal
51.6 GB
Google 4-bit
15.6 GB
Uluka 26B
8.7 GB
12 GBRTX 3060, RTX 4070 SUPERDoes not fitDoes not fitFits
16 GBT4Does not fitDoes not fitFits
24 GBL4, RTX 4090Does not fitFitsFits
32 GBRTX 5090Does not fitFitsFits
48 GBL40SDoes not fitFitsFits
80 GBA100, H100FitsFitsFits

A model counts as fitting if 1.5 GB is left free for it to work.

A much larger model

At 26B the saving is modest: community builds of 11 to 14 GB already fit on the 16 GB GPU. The saving should grow with model size. GLM-5.3-Flash is a 320B model, and its usual 4-bit formats need two to four 96 GB GPUs. A version that fits on one should cut GPU cost by more than half. That is the target for Uluka 320B.

Projected cost per server, per month

Projected monthly cost of running GLM-5.3-Flash in each format, and what Uluka 320B would save
GLM-5.3-Flash versionGPUs needed (96 GB each)Cost per month*Uluka 320B would save per month*
Uluka 320Bour target, about 85 GB1$2,455 projected*–
MXFP4calculated, 170 GB2$6,049$3,594 (59%)
Q3UD-IQ3_XXS, Unsloth, 120 GB2$6,049$3,594 (59%)
NVFP4NVIDIA, 204 GB4$12,098$9,643 (80%)
Q4UD-Q4_K_XL, Unsloth, 200 GB4$12,098$9,643 (80%)

*Projected: Uluka 320B is not built yet. Prices are Amazon Web Services g7e instances, on-demand, US East, September 2026: $3.36 an hour for one GPU, $8.29 for two, $16.57 for four. Versions that need the same number of GPUs cost the same. MXFP4 is calculated from its size, because we found no published build. Speed is not compared.