Framework Desktop · Local AI research brief

Own the
compute.
Route the work.

A practical analysis of the Ryzen AI Max+ 395 in 64/128GB configurations versus the Ryzen AI Max+ PRO 495 with 192GB — for text, coding, image/video generation, and an always-on autonomous agent stack.

Measured where possible Linux-first inference Research updated 1 Oct 2026
The central finding
2× 395/128

At roughly the same capital cost as one 495/192, two 395/128 nodes deliver 256GB total RAM, 80 GPU CUs, 32 CPU cores and ~512GB/s aggregate theoretical memory bandwidth.

Executive thesis

The 495 is a capacity upgrade.
Two 395s are a throughput upgrade.

For interactive single-model work, 192GB can unlock models that simply do not fit in 128GB. For an always-on agent farm, independently scheduled jobs are usually more valuable than one giant address space.
Framework’s own framing: Gorgon Halo retains the same CPU/GPU architectures as Strix Halo; memory bandwidth rises about 6.7%, while the major local-AI change is the additional 64GB of unified memory.
Implication: if your work is parallelizable, an additional 395-class node gives you a second memory bus and a second GPU instead of asking two resident models to share one.
Hardware matrix

Three configurations, two very different buying philosophies.

Ryzen AI Max+ 395
64GB
Value / fast workers
GPURadeon 8060S · 40 CU
MemoryLPDDR5X-8000 · 256-bit
Theoretical BW~256 GB/s
Best fit~27–35B class + quantized media
Ryzen AI Max+ 395
128GB
Local-AI sweet spot
GPURadeon 8060S · 40 CU
MemoryLPDDR5X-8000 · 256-bit
Theoretical BW~256 GB/s
Best fit100–250B sparse MoE + BF16 media
Ryzen AI Max+ PRO 495
192GB
Single-node capacity
GPURadeon 8065S · 40 CU
MemoryLPDDR5X-8533 · 256-bit
Theoretical BW~273 GB/s
Best fit>128GB working sets / dual resident models
+50%

RAM going from 128GB → 192GB.

+6.7%

Theoretical memory-bandwidth increase cited by Framework.

+3.4%

GPU clock: 2.9GHz → 3.0GHz from AMD’s published specs.

≈ same

Core architecture and 40-CU GPU scale; capacity, not a generational compute jump.

Text + coding inference

What actually feels fast — and what merely fits.

64GB → 128GB mostly changes which models fit. A model that already fits does not become faster just because more RAM is installed. Sparse MoE models are especially attractive on this memory-bandwidth-limited architecture.

Measured on 395-class Strix Halo or Framework Estimate / extrapolation Vendor/platform fact
Model / quant395 / 64GB395 / 128GB495 / 192GBRead
Qwen3.6 35B-A3B Q4~55–80 t/s~55–80 t/s~60–85 t/s est.Excellent all-round sparse MoE
Qwen3.8 27B Q4~26.5 t/s measured~26.5 t/s measured~28 t/s est.Dense-ish quality path; MTP matters
Qwen3.8 27B Q8~15.2 t/s measured~15.2 t/s measured~16 t/s est.Higher precision, still interactive
70B dense Q4Fits, ~5 t/s class~5 t/s classSlightly fasterCapacity ≠ pleasant decode speed
GPT-OSS 120B MXFP4Not practical~50–56 t/s measuredSlightly fasterWhy sparse models change the equation
~120B sparse MoE Q4Low quant / tight~18–30 t/s class~20–32 est.Strong “senior local model” tier
DeepSeek V4.1 Flash Q2NoMain weights exceed comfortable single-node fit~14 t/s; + speculation upliftCanonical 192GB use case
MiMo V2.6 FlashNoNo comfortable fit~18 → 16 t/s as context growsLarge single-node model tier
Coding model395 / 64GB395 / 128GB495 / 192GBUse case
Qwen3-Coder 30B-A3B Q4~97–99 t/s~97–99 t/s~100–105 est.Fast terminal / editor agent
Qwen3.6 35B-A3B Q4~60–77 t/s balanced~60–77 t/s~65–82 est.Generalist that also codes well
Qwen3.8 27B Q4~26.5 t/s~26.5 t/s~28 est.Slower dense quality path
120B-class sparse reviewerUsually too tight~18–30 t/s class~20–32 est.Architecture / debugging / review
Huge >128GB reviewerNoNo single-node fitYesWhere the 495 becomes unique
Serving matters: a Strix Halo benchmark of Qwen3-Coder 30B-A3B shows ~90 t/s single-request server throughput and ~173 t/s aggregate at 8 concurrent requests. That makes the 395 unusually interesting as a multi-agent API box, not just a one-user chat machine.

1. Fit is not speed

A 70B dense model can fit comfortably and still decode slowly because most weights must be streamed through memory every token.

2. Active params matter

Sparse MoEs can have very large total parameter counts while activating only a small subset per token — a much better match for the bandwidth profile.

3. Context costs memory

Do not size RAM against GGUF weight size alone. Leave room for OS, runtime, KV cache, context, multimodal projectors and speculative/MTP components.

Comfortable weight-file ceilings: think roughly ~45–52GB on 64GB, ~95–110GB on 128GB, and ~155–170GB on 192GB when you want useful context and system headroom.
Image + video generation

Unified memory removes the VRAM wall. It does not remove the compute wall.

Image generation

SDXLFast / easy fit
FLUX.1 Dev~minute-class
Qwen Image 54GB pipeline~2–2.5 min total in measured workflow

Measured Framework Desktop 395/128 Qwen Image workflow: ~14s denoise + substantial VAE/workflow overhead; ~120–150s total depending on steps and resolution.

Video generation

LTX-2 FP8, ~4 sec clip~12 min
LTX-2 FP8, ~10 sec clip~30 min
Wan 2.2 Lightning, 73f 480p~30 min
Wan 2.2, 40 steps~5 hours
RAM progression for media: 64GB already handles a surprisingly wide set of quantized image/video workflows. 128GB buys higher precision, fewer offload gymnastics, and larger graphs. 192GB mostly expands headroom and multi-model residency; it does not turn a 30-minute render into a 3-minute render.
Economics

At your quoted prices, the comparison gets uncomfortable for the 495.

1 × 495 / 192GB

capacity
≈ €7.7k

One compute domain. Unique advantage: single models / working sets that need more than 128GB.

Total RAM192GB
GPU CUs40
Aggregate memory BW~273 GB/s

2 × 395 / 128GB

throughput
≈ €7.8k

Two independent schedulers, memory buses and GPUs. Better fit for separable agent workloads and concurrent serving.

Total RAM256GB
GPU CUs80
Aggregate memory BW~512 GB/s

Prices above are the approximate euro figures supplied for this analysis; they are not Framework list-price claims. Aggregate specs for two nodes are additive only for independent / distributed workloads — they do not magically become one 256GB coherent memory pool.

Why two boxes can matter more than “2×”
Two models loaded ≠ two models at full speed

One 495 / 192GB

Two resident models still share the same ~273GB/s theoretical memory fabric, the same 40 GPU CUs and the same package power envelope. Concurrency creates contention.

Two 395 / 128GB nodes

Model A can saturate one ~256GB/s memory subsystem while Model B uses the other. The scheduler can prioritize interactive work on one node and batch/background work on the other.

Always-on agent architecture

Treat local models like a private inference cluster — not like chat windows.

Tier 1 · Fast local workers

30–35B sparse MoE. Coding, browsing, extraction, test generation, summaries, monitoring, routine shell work.

Tier 2 · Local reviewer

~100–250B sparse model. Architecture, hard debugging, synthesis, conflict resolution, review and escalation.

Tier 3 · Frontier API

Use selectively for genuinely hard, high-value or enormous-context tasks. The cloud becomes an escalation layer, not the default.

research queue
coding queue
monitoring queue
media queue
Important design choice: agents should exchange artifacts more often than endless prose conversations — task specs, patches, test output, structured JSON, decision records and references. This reduces context inflation and repeated token processing.
Agent serving changes the benchmark

Optimize for useful tokens per hour, not a single chat session’s peak tok/s.

~98 t/s

Measured Qwen3-Coder 30B-A3B Q4 single-stream class on tuned Strix Halo.

~173 t/s

Measured aggregate server throughput at eight concurrent requests in one benchmark route.

2 nodes

Let an interactive node stay responsive while a background node consumes long queues.

L3 cloud

Spend API money only where it changes task success, not on routine localizable work.

24/7 economics

Electricity matters.
Depreciation matters more.

For an always-on worker, buying hardware is economically sensible only if it stays meaningfully utilized. The local-vs-cloud argument becomes much stronger when the system is continuously chewing through useful queued work.

Illustrative 100W continuous average
876 kWh / year / node
0.1 kW × 24 h × 365 days = 876 kWh/year
At €0.25 / kWh≈ €219 / year
At €0.30 / kWh≈ €263 / year
At €0.40 / kWh≈ €350 / year

100W 24/7 is deliberately simple and fairly aggressive. Independent measurements put the 395 Framework Desktop around ~10–12W at idle, ~81–104W average in broad load suites, with substantially higher peaks depending on workload.

Decision framework

Buy for the bottleneck you can prove — not the one you can imagine.

Start here
395 / 128

Strongest first node for this plan. It spans fast 30–35B workers, 100B+ sparse models, serious image/video workflows and enough memory to learn what your real queues look like.

StrategyLocal-first + API escalation
RiskLow capital lock-in
Scale when utilized
+ second node

If queues stay saturated, add another 395-class machine or its successor. This buys real concurrency, fault isolation and another complete memory/GPU subsystem.

StrategyDistributed worker / reviewer
Best forAlways-on agent farm
Choose only if needed
495 / 192

The compelling reason is specific: you need a single >128GB working set, or you value keeping two substantial models resident in one box more than aggregate throughput per euro.

StrategyMaximum single-node fit
Trade-offLower compute / €
Bottom line: for an “always-on agents doing grunt work” vision, the 395/128 is the more rational first purchase. Architect the software as distributed from day one. Let real utilization tell you whether the next €4k should become a second node, a future-generation node, or API spend.
Sources & methodology

Measured data first. Estimates labeled.

Performance varies with runtime, kernel, ROCm/Vulkan stack, quantization, context length, speculative decoding, thermals and model revision. This brief uses current Framework/AMD specifications and public Strix Halo measurements, then clearly marks 495 same-model speeds as estimates unless directly measured.

Prepared 1 October 2026 · Prices used in the decision section are user-supplied approximations. No claim is made that aggregate dual-node memory is coherent/shared memory.