Skip to content
AI Compute Radar
IndexCollected 5 h ago
Field guide — Issue 018Friday · 18 Sep 2026 · Models + compute decision layer

AI Right Now.

The AI ecosystem ships faster than anyone can read. We track model momentum, local hardware fit and compute prices, then turn them into one clear answer.

“What should I actually use today?”

One decision, not 50 tabs.

Today’s editionIssue 018 · Friday · 18 Sep 2026

Since the previous edition · newest reading 2026-09-18 11:51 UTC

  1. Heat mover+6Phi-4 Mini · 24 → 30 · Heat Score2026-09-17 18:15 UTC2026-09-18 11:51 UTC
  2. New on the radar this week1DeepSeek-R12026-09-15 18:15 UTC2026-09-18 06:15 UTC
  3. GPU rent down−6.8%A100 · $0.88 → $0.82 per hour · Vast.ai daily median2026-09-16 23:00 UTC2026-09-17 23:00 UTC
  4. Heat mover+5Gemma 4 12B · 47 → 52 · Heat Score2026-09-17 18:15 UTC2026-09-18 11:51 UTC
  5. Cooling−4Qwen3 32B · 30 → 26 · Heat Score2026-09-17 18:15 UTC2026-09-18 11:51 UTC

Pick of the weekWeek 37, 2026

Gemma 4 12B

Gemma 4 12B is this week's pick by rule: the largest counted Heat Score rise among tracked models that run comfortably on a consumer card, from 45 to 62 since the evening of 7 September. Its trending score is back where it stood in late August, its measured Q4_K_M file is 6.6 GB, and it runs with room to spare on any 16 GB card at 8K context.

Read the report →
  • Qwen3.8 27B · HF trending 608
  • Qwen3.8 Flash-Next · HF trending 226
  • Llama 3.1 8B · HF trending 216
  • GLM-5.3-Flash · HF trending 163
  • RTX 4090 $0.53/h
  • RTX 5090 $0.67/h
  • A100 $0.71/h
  • H100 $3.14/h
  • H200 $4.74/h
  • B200 $8.75/h

00 / Radar view

Every tracked model on one scope.

A contact appears at the center when its Hugging Face repository is created and drifts outward as it ages. Brightness is its Heat Score. Bearing groups models by maker and means nothing else.

Radar scope of every tracked model46 models plotted by age since their Hugging Face repository was created; brightness follows the Heat Score.30 d6 mo1 y2 yQwen3 0.6B · Alibaba Qwen · Heat 67 · repo created 2025-04-27 (509 d ago)Qwen3 14B · Alibaba Qwen · Heat 38 · repo created 2025-04-27 (509 d ago)Qwen3 32B · Alibaba Qwen · Heat 26 · repo created 2025-04-27 (509 d ago)Qwen3 8B · Alibaba Qwen · Heat 79 · repo created 2025-04-27 (509 d ago)Qwen3 8BQwen3-Coder 30B-A3B · Alibaba Qwen · Heat 35 · repo created 2025-07-31 (414 d ago)Qwen3 4B · Alibaba Qwen · Heat 47 · repo created 2025-08-05 (409 d ago)Qwen3-Coder Next · Alibaba Qwen · Heat 60 · repo created 2026-01-30 (231 d ago)Qwen AgentWorld 35B-A3B · Alibaba Qwen · Heat 21 · repo created 2026-06-22 (88 d ago)Qwen3.8 27B · Alibaba Qwen · Heat 83 · repo created 2026-08-05 (44 d ago)Qwen3.8 27BQwen3.8 2.4T-A95B · Alibaba Qwen · Heat 42 · repo created 2026-08-08 (41 d ago)Qwen3.8 Flash-Next · Alibaba Qwen · Heat 79 · repo created 2026-08-24 (25 d ago)Qwen3.8 Flash-NextDeepSeek-R1 · DeepSeek · collecting · repo created 2025-01-20 (606 d ago)DeepSeek-V4-Flash · DeepSeek · Heat 59 · repo created 2026-07-31 (49 d ago)DeepSeek-V4-Pro · DeepSeek · Heat 38 · repo created 2026-08-13 (36 d ago)Gemma 3 27B · Google · Heat 43 · repo created 2025-03-01 (566 d ago)Gemma 4 E4B · Google · Heat 59 · repo created 2026-03-02 (200 d ago)Gemma 4 26B-A4B · Google · Heat 74 · repo created 2026-03-11 (191 d ago)Gemma 4 31B · Google · Heat 80 · repo created 2026-03-11 (191 d ago)Gemma 4 31BGemma 4 12B · Google · Heat 52 · repo created 2026-05-23 (118 d ago)Granite 4.2 30B · IBM · Heat 42 · repo created 2026-08-07 (42 d ago)Granite 4.2 3B · IBM · Heat 48 · repo created 2026-08-07 (42 d ago)Granite 4.2 8B · IBM · Heat 46 · repo created 2026-08-07 (42 d ago)Ling 3.0 Flash · InclusionAI · Heat 23 · repo created 2026-08-02 (47 d ago)Ling 3.0 Tiny · InclusionAI · Heat 25 · repo created 2026-08-10 (39 d ago)KAT-Coder V2.5 · Kwaipilot · Heat 35 · repo created 2026-07-23 (57 d ago)LFM2.5 1.2B · Liquid AI · Heat 18 · repo created 2026-01-06 (255 d ago)LFM2.5 8B-A1B · Liquid AI · Heat 32 · repo created 2026-05-28 (113 d ago)LFM2.5 2.6B · Liquid AI · Heat 28 · repo created 2026-07-28 (52 d ago)Llama 3.1 8B · Meta · Heat 83 · repo created 2024-07-18 (792 d ago)Llama 3.1 8BLlama 3.2 3B · Meta · Heat 77 · repo created 2024-09-18 (730 d ago)Llama 3.3 70B · Meta · Heat 59 · repo created 2024-11-26 (661 d ago)Phi-4 Mini · Microsoft · Heat 30 · repo created 2025-02-19 (576 d ago)Mistral Small 3.2 · Mistral AI · Heat 37 · repo created 2025-06-19 (456 d ago)Devstral Small 2 24B · Mistral AI · Heat 31 · repo created 2025-11-28 (294 d ago)Nemotron Nano 9B v2 · NVIDIA · Heat 16 · repo created 2025-08-12 (402 d ago)gpt-oss 120B · OpenAI · Heat 68 · repo created 2025-08-04 (410 d ago)gpt-oss 20B · OpenAI · Heat 70 · repo created 2025-08-04 (410 d ago)Ornith 1.5 35B-A3B · Ornith AI · Heat 72 · repo created 2026-08-18 (31 d ago)Ornith 1.5 397B · Ornith AI · Heat 42 · repo created 2026-08-18 (31 d ago)Ornith 1.5 9B · Ornith AI · Heat 69 · repo created 2026-08-18 (31 d ago)Hy3 · Tencent · Heat 35 · repo created 2026-07-02 (78 d ago)GLM-4.5-Air · Zhipu AI · Heat 29 · repo created 2025-07-20 (425 d ago)GLM-4.7-Flash · Zhipu AI · Heat 40 · repo created 2026-01-19 (242 d ago)GLM-5.2 · Zhipu AI · Heat 67 · repo created 2026-06-16 (94 d ago)GLM-5.3 · Zhipu AI · Heat 70 · repo created 2026-08-25 (24 d ago)GLM-5.3-Flash · Zhipu AI · Heat 76 · repo created 2026-08-25 (24 d ago)

How to read the scope

  • Rings, from the center: 30 days, 6 months, 1 year, 2 years since the repository was created. The rim is three years and beyond.
  • Brightness: Heat Score 0 to 100. A dim outline is a model still collecting its first week of history.
  • Pulsing ring: repository created within the last 30 days.
  • Bearing: maker sector, nothing more. Blips open the model page.

Newest contacts

  1. GLM-5.3Zhipu AI · repo created 2026-08-25 · 24 d agoHeat 70
  2. GLM-5.3-FlashZhipu AI · repo created 2026-08-25 · 24 d agoHeat 76
  3. Qwen3.8 Flash-NextAlibaba Qwen · repo created 2026-08-24 · 25 d agoHeat 79
  4. Ornith 1.5 35B-A3BOrnith AI · repo created 2026-08-18 · 31 d agoHeat 72
  5. Ornith 1.5 397BOrnith AI · repo created 2026-08-18 · 31 d agoHeat 42

Strongest returns

  1. Qwen3.8 27BAlibaba Qwen · 44 d ago83
  2. Llama 3.1 8BMeta · 792 d ago83
  3. Gemma 4 31BGoogle · 191 d ago80

Ages come from Hugging Face repository creation dates; the Heat Score from measured signals. How the Heat Score is measured

01 / Model market

What is winning

Models ranked by Heat Score, with the measured signals behind it
RankModelHeat0–100 · compositeDownloadsHugging Face
1Qwen3.8 27BAlibaba Qwen · OpenGeneral · Coding · Agents · ResearchHugging Face ↗ (External link)GGUF ↗ (External link)Trending: 608$/M in: $0.217.5M+1.8% rising
2Llama 3.1 8BMeta · OpenGeneral · CodingHugging Face ↗ (External link)GGUF ↗ (External link)Trending: 2165.9M+4.6% rising
3Gemma 4 31BGoogle · OpenGeneral · ResearchHugging Face ↗ (External link)GGUF ↗ (External link)Trending: 50$/M in: $0.099.1M+4.2% rising
4Qwen3.8 Flash-NextAlibaba Qwen · OpenGeneral · Coding · AgentsHugging Face ↗ (External link)GGUF ↗ (External link)Trending: 226$/M in: $0.15706K+25.2% rising
5Qwen3 8BAlibaba Qwen · OpenGeneral · Coding · AgentsHugging Face ↗ (External link)GGUF ↗ (External link)Trending: 62.5$/M in: $0.1213.1M+1.4% rising
6Llama 3.2 3BMeta · OpenGeneralHugging Face ↗ (External link)GGUF ↗ (External link)Trending: 29$/M in: $0.051.8M+17.4% rising
7GLM-5.3-FlashBreakoutZhipu AI · OpenGeneral · Coding · AgentsHugging Face ↗ (External link)GGUF ↗ (External link)Trending: 163$/M in: $0.092.4M+139.1% rising
8Gemma 4 26B-A4BGoogle · OpenGeneral · ResearchHugging Face ↗ (External link)GGUF ↗ (External link)Trending: 18$/M in: $0.099.6M+8.4% rising
9Ornith 1.5 35B-A3BOrnith AI · OpenGeneral · Coding · AgentsHugging Face ↗ (External link)GGUF ↗ (External link)Trending: 43394K+25.0% rising
10GLM-5.3CoolingZhipu AI · OpenGeneral · Agents · ResearchHugging Face ↗ (External link)GGUF ↗ (External link)Trending: 47$/M in: $1.40838K+51.9% rising
11gpt-oss 20BOpenAI · OpenGeneral · Agents · ResearchHugging Face ↗ (External link)GGUF ↗ (External link)Trending: 26$/M in: $0.036.7M+1.3% rising
12Ornith 1.5 9BOrnith AI · OpenGeneral · CodingHugging Face ↗ (External link)GGUF ↗ (External link)Trending: 21506K+35.4% rising

Top 12 of 46 tracked models, ranked by Heat Score; models still collecting history follow, ranked by trending.

Heat is a composite of measured signals only: each component is the model's percentile among models with a full week of history — trending level (35%), 7-day download growth (30%), 7-day trending change (15%), 30-day downloads (15%) and Hub likes (5%). Missing components renormalize the weights and lower the shown confidence; nothing is guessed. Trending and downloads come from the Hugging Face Hub — the download counter is a rolling 30-day window, not unique users — and input pricing from OpenRouter (CC BY 4.0). The full formula, thresholds and flag rules are on the methodology page. How Heat is computed → Every tracked model has its own page →

02 / Compute ladder

Own, Spark or rent?

Models need machines. The right machine depends on how often you push past your own hardware — not on the biggest number on a spec sheet.

Path A — Existing PC

Use what you own.

Most people underestimate their current GPU. A 12–32 GB card runs excellent quantized models every day, free of charge.

  • + 12–32 GB VRAM class
  • + Ollama / LM Studio ready
  • + Zero extra cost

Path B — Personal AI box

Own the middle.

DGX Spark and GB10-class systems put 128 GB of unified memory on your desk. Rational when large local models are your daily routine.

  • + 128 GB unified memory
  • + 273 GB/s — capacity, not speed
  • + Private by default

Path C — Elastic compute

Rent the spike.

Occasional heavy job? Rent an H100 for the afternoon instead of buying hardware that idles the rest of the year.

  • + H100 / A100 / B200 class
  • + Pay per hour
  • + No commitment
Median $/h
  • RTX 4090$0.53steady
  • RTX 5090$0.67rising
  • A100$0.71steady
  • H100$3.14rising
  • H200$4.74steady
  • B200$8.75rising

03 / My setup

Your machine, your answer

EstimatedEstimated: computed from our curated model and hardware catalog — not a live reading.

Four questions. One recommendation. No benchmark knowledge required.

Runs locally in your browser — nothing is sent to any server.
Saved only in this browser (localStorage), never sent to us. Untick to forget.

Local first. — Fit score 73

Local first.

Main model
GLM-4.7-Flash · 31.2B MoE · Q4_K_M
Compute
RTX 4090 · 24 GB VRAM
Burst option
Vast.ai when the job outgrows you
Estimated memory
~18.9 GB

Why this route

  • + Weights ~17.1 GB (Q4_K_M), context ~0.4 GB at 8K tokens, runtime ~1.4 GB.
  • + Owning the inference beats paying per token at this frequency

Hugging Face ↗ (External link)

Estimated from quantized model file sizes plus runtime overhead and a safety margin. Real usage varies with context length.

Tap the bear to wake it

From the channel

Short answers, on YouTube

The same numbers as on this site in videos under a minute: does this model fit that card, and why. Measured file sizes, computed context memory, no speed claims.

04 / Hardware radar

Models need machines.

EstimatedEstimated: computed from our curated model and hardware catalog — not a live reading.

Modelsneedmachines.

We track the machines that matter for local AI — consumer GPUs, personal AI systems and the rentable heavy metal.

Hardware watchlist

DeviceMemorySignal
RTX 4070Entry local12 GBSteadysteady
RTX 3060 12GBBudget classic12 GBHuge install basesteady
RTX 4080Solid local16 GBSteadysteady
RTX 5080Current 16 GB16 GBCurrent genrising
RTX 4060 Ti 16GBBudget 16 GB16 GBValue picksteady
RTX 5060 Ti 16GBCurrent budget 16 GB16 GBCurrent genrising
RTX 5070 TiFast 16 GB16 GBCurrent genrising
RTX 3090Used-market value24 GBDemand risingrising
RTX 4090Local workhorse24 GBDemand risingrising
RTX 5090Top consumer32 GBDemand risingrising
Mac mini M4 Pro 64GBApple unified 64 GB64 GB UApple Siliconrising
Mac Studio M4 Max 64GBApple unified 64 GB64 GB UApple Siliconrising
Mac Studio M3 Ultra 96GBApple unified 96 GB96 GB UApple Siliconrising
NVIDIA DGX SparkPersonal AI system128 GB UNew categoryrising
A100 80 GB (rental)Rental workhorse80 GB$/h fallingfalling
H100 80 GB (rental)Rental flagship80 GB$/h trackedrising
B200 180 GB (rental)Rental frontier180 GB$/h fallingfalling

Radar Card of the day

RTX 4090

In 2022 this machine shipped with 24 GB. Today 22 of 46 tracked models run on it at 8K context.

Open the card →All cards →