Skip to content
logo

AIFinitee

All things AI & LLM

  • Home
  • Articles
  • Blog
Subscribe
Top Stories
MiniMax H3: Your guide to proper prompting for better videos
GLM 5.2 at 4-Bit With MTP: The Easiest Way In
Running DeepSeek V4 Flash Q2 on Strix Halo 128GB
ComfyUI on AMD Ryzen AI MAX: 96 GB Unified Memory vs 16 GB NVIDIA
Speed vs. Smarts: When Bigger Models Win for Local AI Coding
Linux vs Windows for LLM Inference: More Tokens/Second on Linux
Strix Halo Cluster
Posted inMiniMax

MiniMax H3: Your guide to proper prompting for better videos

As much as cloud remains a dominant option for the casual consumer, it is worth to remember that AI (LLM or units of compute) are also units of Intellect. Ownership…
Continue Reading
Posted by
Posted inApple Silicon

GLM 5.2 at 4-Bit With MTP: The Easiest Way In

And the 512GB Mac Studio that Apple just stopped selling. Apple discontinued the 512GB Mac Studio. Not the chip — the M3 Ultra is still on the shelf — but…
Continue Reading
Posted by
Posted inAMD Distributed Inference

Running DeepSeek V4 Flash Q2 on Strix Halo 128GB

AIfinitee belives local compute as an alternative to cloud and that the most capable open models should be runnable on hardware you actually own. This post is the recipe for…
Continue Reading
Posted by
Posted inAMD Distributed Inference

ComfyUI on AMD Ryzen AI MAX: 96 GB Unified Memory vs 16 GB NVIDIA

At AIFinitee, we have spent months benchmarking LLM inference on our dual-node AMD Ryzen AI MAX 395+ cluster. We have measured tokens per second across MiniMax-M2 and Qwen3.5-397B. We have…
Continue Reading
Posted by
Posted inAMD Distributed Inference

Speed vs. Smarts: When Bigger Models Win for Local AI Coding

At AIFinitee, we've spent months chasing tokens per second. Our two-node AMD Ryzen AI MAX cluster hits 17-20 tok/s with MiniMax-M2. Our Linux-vs-Windows benchmarks showed how your OS quietly taxes…
Continue Reading
Posted by
Posted inAMD Distributed Inference

Linux vs Windows for LLM Inference: More Tokens/Second on Linux

Same GPU, different OS = different performance. See why Linux delivers 5-30% more tokens/sec than Windows across NVIDIA, AMD, and Intel GPUs.
Continue Reading
Posted by
You May Have Missed
Posted inMiniMax

MiniMax H3: Your guide to proper prompting for better videos

Posted by
Posted inApple Silicon

GLM 5.2 at 4-Bit With MTP: The Easiest Way In

Posted by
Posted inAMD Distributed Inference

Running DeepSeek V4 Flash Q2 on Strix Halo 128GB

Posted by
Posted inAMD Distributed Inference

ComfyUI on AMD Ryzen AI MAX: 96 GB Unified Memory vs 16 GB NVIDIA

Posted by
Copyright 2026 — AIFinitee. All rights reserved. Bloghash WordPress Theme
Scroll to Top