Rendered at 13:54:44 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
steinvakt2 18 minutes ago [-]
Any way to use this for speech-to-text? To gain faster whisper inference for instance, without losing accuracy?
lxe 20 hours ago [-]
On my local inference box I have a perpetual codex thread open in my llama.cpp checkout that I periodically ask to take a look at currently pending llama.cpp PRs, do some research on latest MTP, Dflash and other prediction or attention optimizations, do research on the latest model quants and finetunes, take a look at localLlama Reddit threads and just do essentially a sweep of the frontier.
Then it rebuilds latest llama.cpp, grabs the PRs it finds relevant to test against, and then it performs a benchmark and finalizes the upgrade and verifies what model, variant, or even a separate finetune that we should be running.
Occasionally, it performs its own optimizations and commits, which then gets superseded by pull requests and merged code that essentially validates the model's own optimization directionality.
anerli 18 hours ago [-]
Yeah we heavily leverage coding agents for optimizing our kernels. Since it's highly verifiable and takes time to measure we often leave multiple running and improving performance on different model architectures.
Definitely still helps to reference relevant academic work as well, or even just encouraging the agent to make bigger structural leaps, otherwise it will often get stuck working on low impact micro-optimizations.
bredren 11 hours ago [-]
What kinds of prompting do you use to encourage structural leaps in your agents?
kmike84 20 hours ago [-]
This seems to be a good idea. However, beating llama.cpp on speed is a low bar :)
I found it to be a good baseline, but at least on Mac there was always something way faster, and/or with better memory requirements - like you said, ds4, omlx, mtplx, etc. It seems if you use local LLMs for real, there is very little reason not to use one of the more optimized engines.
3 main failure modes I observed in the engines:
* Not using best available spec decoding
* Using too much VRAM for KV cache (e.g. KV cache used to take almost nothing in ds4, but huge amount of VRAM on unsloth/llama.cpp for deepseek models)
* Degraded performance at large context sizes - benchmarks at 4K or 32K are awesome, but at realistic 100-200K it's slower than some stupid baseline
anerli 19 hours ago [-]
Yeah these are all things that we directly tackle!
Spec decoding: Models in our catalog come assigned with an assigned drafter model for speculative decoding based on the best known method and model available for that target model (support DFlash, DSpark, and DFlash2).
Using too much memory for KV cache: We use a TurboQuant-inspired quantization of KV cache to 8-bit keys and 4-bit values. This drops KV memory usage by over half and also speeds up decode. Based on long context quality benchmarking we've done it does not seem to negatively impact retrieval or coherence over long context.
Large context sizes: our KV quantization helps a lot for this, and we focus our optimizations on specifically longer-context requests since that's what most agent inference actually looks like.
kmike84 17 hours ago [-]
I think the specific issue I had was due to lack of proper support for ds4 compressed KV cache, not about KV cache quantization. It was like 50GB instead of 5GB for context, and it wasn't fixed for weeks (I haven't checked if it's fixed now - hopefully it is).
Quantization is another thing. There are so many engines launched with claims about speed, but in many cases it's optimizing specific lower-quality quants. When you have enough resources, you usually want something like W8A16 + full precision KV cache working as fast as possible, not yet another W8A8 or W4A16.
In general, it seems new models are released so fast now - engines don't always have time to really polish the implementation before the next model is released
sheephess44 11 hours ago [-]
[dead]
skohan 19 hours ago [-]
Do you have anything published on the quality benchmarking using your caching strategy?
anerli 19 hours ago [-]
We have a retrieval benchmark based on RULER which we've been using to ensure that the model maintains complete awareness of the full context window.
and the results are what i kinda expected to begin with, this adds next to nothing? - also the repo was pivoted from a playwright sub assembly to this not long ago - so my conclusion - THIS MIGHT be worth some watching if you have a model where no-one!, has optimized it at all - and where it does not use anything native to your platform.
results (averages)
VIA rapid-mlx :8902 (MLX)
Decode: 175 tok/s
TTFT: 64 ms
prefill (~760 tok cold): 594 ms
Magnitude 0.2.1 (GGUF/llama.cpp+Rust)
Decode: 161 tok/s
TTFT: 111 ms
prefill (~760 tok cold): 669 ms
anerli 6 hours ago [-]
Hi, M5+ Macs have a known optimization gap since we do not fully utilize the Metal 4 matmul operations in our kernels yet - so should be able to do much better on this specific comparison soon!
pbronez 2 hours ago [-]
Rapid-MLX was my first thought too. It’s optimized for self-hosted agents on Apple silicon. It’s my current choice for self-hosted models.
On a Mac the baseline I'd want is MLX, not llama.cpp. llama.cpp isn't the fast path on Apple Silicon for most models people run locally, so a speedup over llama.cpp could still be slower than mlx_lm. Do you have that number?
Second, more important for agents: decode speed is rarely what hurts. It's resending the same system prompt plus tool schemas every turn. Does self-optimizing cover prefix cache reuse across requests, or is it kernel and layout tuning only?
And is the endpoint OpenAI-compatible? My runtime already talks to llama.cpp and LM Studio through that wrapper, so drop-in is the difference between trying it tonight and not.
bythreads 7 hours ago [-]
i think this is a wash at best, but might be ok if you have a very specific model that is completely unoptimized - but these days you could just point your astra level llm at it and say "make this faster"
anerli 6 hours ago [-]
Yes MLX is generally a better comparison point overall for Apple, planning on releasing a benchmark for that soon. However against the MLX-based engines we've compared with so far Magnitude will continue to have an edge, especially for decode kernels.
Prefix cache is re-used with a prefix tree structure for maximal re-use across sessions sharing prompts.
The endpoint is standard OpenAI compatible chat completions.
kmike84 16 hours ago [-]
How accurate are speed estimates in the UI? I'm asking because for Qwen 3.8 (Q8) the speed numbers cited in the UI look quite poor:
262K number is ok, but for lower context sizes (<128K) it's about 2x slower than the numbers I'm getting from real mtplx sessions for qwen3.8 q8 (mac m5 max).
Is it a lack of optimizations, or incorrect numbers, or a benchmark artifact (e.g. something which is harder for spec decoding than usual agentic sessions)?
anerli 15 hours ago [-]
These numbers are just estimates based on your hardware and may differ from actual performance. It's hard to get an accurate measurement until it's actually downloaded and running. They also don't account for gains from speculative decoding. Working on changes to make this more clear.
For m5 - there may be some issue with the Metal 4 matmul hardware utilization that could be causing this to be behind here. Will look into this.
francisjp 15 hours ago [-]
To OP: great work on the release! I am generally interested in this kind of optimization work.
Related to the post above: Similar results here M5 Max running Qwen3.8 UD-Q6-K-XL with zlab’s Dflash2 as the drafter.
Both prefill and decode are roughly 2x faster when served from llama.cpp (b10853 or newer) than magnitude 0.2.1.
anerli 15 hours ago [-]
Thanks for pointing this out. I think we have a gap here where we may not be fully leveraging the new matmul operations available on M5+ chips, so will work that into our kernels soon and benchmark on M5 hardware.
francisjp 15 hours ago [-]
Sure thing, happy to share. That potential root cause makes sense. I bet magnitude will close the prefill gap then.
llama.cpp had that same matmul gap (pre-fill operations) for the M5/A19 and newer silicon until that sha mentioned above.
mncharity 19 hours ago [-]
Fwiw, top of my own pain-point list (I suppose given the first item, that's a pun) includes:
External/policy-based throttling for temperature control. Unthrottled, my laptop bottom goes skin-burn hot. But fixed compute caps can have non-linearly dreadful performance impacts in particular cases. Plan is a runtime knob, to replace manual limits-kludgery.
I'll use models which barely fit in VRAM+RAM, and are order-1 tok/s slow. So tool call step overhead can be painful - a world where `ls` costs tens of seconds. Plan is blending harness plugins with inference loop, for "no, don't stop - I already have the call result for you - just keep going" (and also some logit games).
anerli 19 hours ago [-]
Temperature control is a good callout! Feel free to open an issue for any features that you'd like to see in Magnitude
I have two NVIDIA GPUs (16GB+16GB) here, and it detects them each twice (says I have 4 GPUs). But then, it says most models are too big (anything >8GB?) and seems to run only on one GPU (5070ti).
Unfortunately even with my 5070ti, llama.cpp seems to be about 20-30% faster at decode, running as:
set CUDA_VISIBLE_DEVICES=0
build\bin\Release\llama-server -hf google/gemma-4-12B-it-qat-q4_0-gguf -ngl 99 --no-mmproj-offload -mg 0 -c 262144 -fa on --host 0.0.0.0
anerli 19 hours ago [-]
Thanks for reporting the issue.
Currently we don't support multi-GPU setups, that is on our near-term roadmap. It saying the model is too big for that GPU might be a bug - would you be willing to open a github issue with more detail on your setup? https://github.com/magnitudedev/magnitude/issues
As for performance, there may be some variability still depending on the model and backend. We have room for improvement for various setups that we are closing as we work out some details with our kernels and tuning system, so appreciate the data point and will look into that combination.
Congratulations on the launch, it looks like an impressive product and tool!
Q: From my (very, very limited!) understanding, I’m under the impression that part of the “inference engine inertia” is that model- or at least architecture-specific code is required for most, if not each new open-weight model coming out.
Assuming I got that right, do you plan on supporting everything vLLM/llama.cpp can do, such that Magnitude becomes a drop-in replacement for as many (economically/pareto-viable) models as possible, or do you want to focus on the best possible support for only a select few models/classes of models?
anerli 19 hours ago [-]
Yeah, generally being able to focus on specific architectures lets you optimize better for those. However models of the same family (for example Qwen 3.5/3.6/ some 3.8 models) share the same architecture, so you only need to optimize once and new models can use the same kernels. There's also shared algorithms and kernels that can be optimized once and used across different families, so it's a bit nuanced.
We plan to support any model architecture that we believe is somewhere along or close to the pareto frontier. There's some model families that are outdated or more niche that we don't necessarily want to put our focus into.
msdz 18 hours ago [-]
Makes much sense and is about what I expected for a project like yours, thanks for the reply.
happybox2016 15 hours ago [-]
2x llama.cpp" on what, an M3 Max? llama.cpp's metal kernels already saturate memory bandwidth. Real agent bottleneck isn't single-stream tok/s — it's KV cache for 5+ concurrent 128k contexts on 24GB VRAM. Who's actually running multi-agent locally? A) Single session only B) 2-3 agents C) 5+ agents D) Gave up,
williamse 8 hours ago [-]
B, 2-3 agents. The KV cache framing is the right one. Single-stream tok/s is what shows up in benchmarks but it's not what actually hurts when agents are sleeping between tool calls and waking up needing their full context. The question I'd want answered about an engine like this is how it handles partially-cold contexts, because agent sessions aren't uniform sustained reads, they're bursty and interleaved.
anerli 14 hours ago [-]
On our benchmarks we approach 2x decode speeds on a variety of Mac hardware (tested most on M4 Pro and Max).
llama.cpp does not saturate memory bandwidth for single-stream tok/s, and for long context and batching, our quantized KV and associated decode kernels allow us to reduce the effective bandwidth needed, and surpass llama.cpp significantly in decode speeds.
larodi 7 hours ago [-]
Everyone focusing on the specs side, but how is this enterprise going to make money, given it is a YC cohort company?
mrtsepelev 4 hours ago [-]
Congrats on launch! Tried it on the gemma-4-26b-qat-4bit model. Was indeed faster on token generation then on oMLX (82.8 tok/s vs 76.5 tok/s), but the prefill time was ~2.6x slower (709 tok/s vs 1843 tok/s). Don’t use any acceleration on the oMLX. Macbook M5 Pro, 48 gb
nateb2022 20 hours ago [-]
Any source on the benchmarks/methodology besides the image? There's a ton of variance possible in llama.cpp's performance depending on how it was configured. I'd also like to see benchmarks against MLX.
anerli 20 hours ago [-]
The benchmark we cited here is a simple prose-repetition task. We put the content of Moby Dick up to 64k context in the request, and then ask it to repeat the last section.
For llama.cpp, we try to make the comparison as fair as possible by using similar settings. No speculative decoding, default prefill batch sizes, flash attention on.
We tried also quantizing the KV cache to 8-bit keys and 4-bit values like we do in Magnitude, but this bombed decode speed for llama.cpp in our testing. Since it seems llama.cpp did not optimize that path, we used 16-bit KV instead.
Compared to MLX - we've done some rough benchmarking and we are outperforming any of the MLX-based engines we've compared to so far. Going to do more in depth benchmarking and release it soon.
karlkloss 2 hours ago [-]
It couldn't detect my hardware at all, Win11 Intel Core Ultra 7 notebook with Nvidia RTX PRO 1000.
thoughtpeddler 9 hours ago [-]
For those running local models on macOS with Apple Silicon, how does Magnitude differ from what Apple's own first-party Core AI now does during its "specialization" procedure, wherein it performs some kind of "model optimization and conversion" (vis-a-vis AOT compilation) into a Core AI "compiled model" file that is optimized for Apple Silicon (leveraging custom Metal 4 kernels, or so Apple says), per WWDC labs from this year that discuss this? [0]
Core AI is a way to ship a model inside an app rather than an inference engine you can point agents to, and it's not meant for agent workloads.
thoughtpeddler 7 hours ago [-]
Understood that's the general use-case, but couldn't one use it 'single-purpose' to ship a model as an inference engine (to point agents to)?
aitoolcrux 8 hours ago [-]
Self-optimizing inference is one of those problems where marginal gains compound—small improvements in batching, KV-cache reuse, and routing across thousands of requests add up fast.
The hard part isn't the optimization itself—it's measuring whether a change actually helps across the long tail of request patterns. Most inference benchmarks show huge gains on popular workloads, but production traffic has a fat tail where naive optimizations hurt latency.
Interested in how you handle regression detection when the optimizer changes between requests.
c7b 18 hours ago [-]
Cool idea! Do you happen to have benchmarks for Strix Halo (AMD Ryzen AI Max+ 395)? I take it that Qwen3.8-Flash-Next is not supported?
And a more general question: does your engine detect and optimize for custom setups like multiple (possibly different) GPUs, eGPUs,...? Because if all you have is a stock major system like a Mac or DGX Spark, that's all you're going to care about, and there are a lot of highly optimized single-hardware engines out there that will be hard to beat in the long run. Something that automatically adapts to custom systems that don't have their own subreddits could really fill a gap.
anerli 18 hours ago [-]
No specific benchmarks for Strix Halo yet but planning to release more results for different hardware and models soon!
Qwen3.8-Flash-Next support will also be added very soon.
Taking full advantage of all the hardware on your machine in the most performant way possible is the overall goal of the inference engine. This includes a lot of what you're describing. We want to map out the full hardware topology of your system (one or more GPUs, CPU, memory), and compile a combination of kernels to serve a given model optimally across that stack, allocating different parts of the workload wherever it fits best.
Currently we're writing tunable kernels that optimize themselves for one device, but we're working on a kernel compiler that will be able to compile and distribute kernels across any number of devices in a system.
lin7c 11 hours ago [-]
One thing I'd want to see in the evals is per-turn latency across a full agent trajectory, not just end-to-end time. In my experience the workload flips mid-run: early turns are prefill-heavy (big system prompt, tool schemas), late turns are short decodes against a huge KV cache, so a config that's optimal for turn one can be badly wrong by turn thirty. The self-tuning idea is interesting, but I'm curious whether the tuning happens per-request or per-trajectory. With prefix-cached tool schemas the win should compound; without it you're re-solving the same optimization problem every call.
singh_abinashi 15 hours ago [-]
Curious how the evals for this work on real agent workloads versus synthetic benchmarks. In my experience, agent cost and latency profiles change a lot once there's a tool-use loop involved, because the token distribution gets much burstier than a single prompt. Did you evaluate on multi-step tool-calling traces, or mostly single-turn?
thoughtpeddler 9 hours ago [-]
This just makes me think about optimizing model weights to run as 'close to the metal as possible' (i.e. within or 'just above' a UEFI boot environment, like the NightRun project), so that there isn't any OS-level overhead either. If we're optimizing, let's optimize! Curious though, maybe the OS doesn't impose much of a burden here? Open to hearing what other tinkerers think...
anerli 9 hours ago [-]
Once you get down to the level of writing GPU kernels the OS isn't very much in the way, aside from some host-side memory logistics.
So it's more about making those GPU kernels perform the required memory move and arithmetic operations as close to the theoretical optimum as possible.
theParadox42 7 hours ago [-]
I remember when this company was just doing agentic playwright style browser interactions
From the looks of it, this seems focused on datacenter/batch inference, and doesn't tune its kernels to the specific hardware and workload where inference is being run like Magnitude does.
Magnitude is optimized for maximum single-session performance and memory efficiency - so we should be more performant for local inference use cases.
hypercube33 19 hours ago [-]
From your description looks like this isn't for AMD or Strix Halo at all? Also one of the things I'm not sure of but definitely plays a huge factor is the variant of the model you download - how does this help select the fastest version for your specific hardware / context size?
anerli 19 hours ago [-]
We support Vulkan as well, we just didn't mention it in the benchmark. When AMD or Strix Halo is detected the engine will use Vulkan.
Regarding model variants - our catalog includes different quantizations, and automatically assesses these against your hardware to determine which ones will fit in your memory and how fast they will run. This lets you pick a model to download based on your desired speed/intelligence tradeoff.
skohan 19 hours ago [-]
Do you have any plans to support ROCm?
anerli 19 hours ago [-]
We are actively benchmarking our Vulkan kernels to ROCm implementations in other engines to ensure that we can reach the performance ceiling with them. Vulkan is much more portable and also works on non-AMD hardware even though it can be more awkward to write kernels for. If we find that Vulkan is not sufficient for reaching the same performance as ROCm, we'll consider adding it as a backend
cedricd 18 hours ago [-]
Looks interesting! Is there any way to skip or speed up the 'Assessing Models' step? I'm unable to download anything because it's been taking forever. I'm sure you could apply some quick heuristics or do a lookup or something to filter models. Or trust the user a bit more -- I already know which models fit on my machine. As it stands I'm stuck at that step and can't use the app.
Maybe have it run silently in the background and assess on demand when a user selects / attempts to download a model. It's not quite clear why all need to be assessed before I can download the first model to try.
anerli 18 hours ago [-]
Thanks for reporting this issue - assessing is not supposed to take more than a minute or so. This is not strictly necessary but filters out models that don't fit in memory and gives speed estimates. This should ideally be a very short step so skipping hopefully wouldn't feel necessary if we patch this.
Could you share your hardware and OS details to help us identify what might be the issue here?
if fully custom compiler would find best settings for given setup, upload the setup to mothership and allow new peers to download it as good starting point.
While it's great to see tok/s go up as high as possible, I think it's important to consider the actual usability of these models when you quantize down to something like 2-bit. From what we've tested it seems like going below 4-bit quickly leads to serious issues with thinking, tool calls, and overall model coherence.
Our plan to enable running bigger models on less GPU memory in a way that'll remain productive is expert streaming. This will let you offload experts for MoE models to RAM or disk, and load them when needed. This can have some performance tradeoff, but is lossless.
MaxikCZ 7 hours ago [-]
I totally get that, Its just, regardless of how bit-quantized it is, they are still pulling 40t/s from model sitting mostly in RAM instead of VRAM. If I understand correctly they split the network parts very deliberately between VRAM and RAM, and I wonder if your program, of which main feature is "get most of your hardware" is capable of similar feats, or if that performace is still locked for those willing to spend days experimenting manually.
digitaltrees 13 hours ago [-]
Do you support splitting models across devices so larger models can run on clusters?
I am building propelcompute.com an open router for private hardware and experimented with exo labs to run large models on for Mac studios and plan on doing the same with nvidia and amd. Id love to integrate your inference engine into the system but built gpu is critical.
anerli 12 hours ago [-]
The goal of the inference engine is to make the best use of whatever hardware you have to run models performantly and let people run bigger models. At first this will include using all the hardware on a given machine optimally. Eventually we also want to support interconnect between multiple machines to enable running bigger models!
digitaltrees 11 hours ago [-]
Let me know if youre interested in a collaboration then. I am working on a custom mlx sharding system.
chzblck 11 hours ago [-]
Sounds interesting would love to test it out but here's what I got when I first tried to get some models downloaded.
On a 64gb Ram and 5080 machine the biggest model suggested was Qwen 9B
can hit 90+ tps on the MoE 35b but mag thinks it wont fit.
anerli 8 hours ago [-]
Yeah that doesn't sound quite right. Given that the 5080 has 16 GB of VRAM I would have expected a few more options, for example Gemma 12B, to at least show up as available. Do any bigger models show as available or was that specifically the recommendation?
The MoE 35b might be tight though unless you were to go below 4-bit. Could you share the quant you used when you ran this on that 5080 before? Our catalog only contains models down to 4-bit because we find that thinking, tool calling, and overall capabilities start to suffer at lower fidelity.
Feel free also to create a GitHub issue with more details and we can take a closer look.
kenzic 20 hours ago [-]
How long does tuning take (on an M3 MacBook Pro for example)?
anerli 20 hours ago [-]
Tuning is a one-time process that takes around ~1 minute whenever you download a new model. This is generally enough time to tune all the kernels' parameters to the point where tuning any longer asymptotes. Time can vary a little based on the hardware though.
kenzic 20 hours ago [-]
Wow, that's impressive.
amirhesham 20 hours ago [-]
Oh this is so cool. Curious about the business model, too.
anerli 18 hours ago [-]
As mentioned here https://news.ycombinator.com/item?id=49912327 we'll eventually build an inference cloud for hybrid workloads. For now though, we're focused on making the inference engine great!
loclol101 9 hours ago [-]
Will it support multi-agent setup across heterogeneous devices (macbook pro, RTX5090, mac mini, etc)?
anerli 8 hours ago [-]
In the short term we will enable using multiple devices within one machine efficiently and distributing the workload of a model between them. If you mean running the same model distributed across different devices, yes we do have plans to enable that. You should eventually be able to run larger models across machines with capable hardware as long as you can provide a fast enough connection between them.
NKosmatos 4 hours ago [-]
Nice one!
Tried it and unfortunatelly there are no small models that can fit my 16GB RAM or GTX1650 4GB GPU. Yeah, I know that this configuration is not meant to be used for AI/LLM work, but it would be good to provide support for some smaller models so that us plebeians can also play a bit with what you techbros are used to ;-)
There are many small/very small models out there and I'm sure you could add a couple just for playing around and experimenting.
paulgerhardt 19 hours ago [-]
Trying to run this but keep hitting bugs. Can you open up issues reporting on your repo?
anerli 19 hours ago [-]
Issues should already be open for anyone to report!
Let me know if you keep running into problems for some reason
sgtwompwomp 20 hours ago [-]
This is dope, is this kind of like Wafer.ai but for local models? As in a coding agent optimizes the kernels so the local model runs continuously better? Cause that is compelling if so. If it’s more simple that’s cool too
anerli 19 hours ago [-]
I would say the overall idea of trying to achieve performant inference for agent workloads is the strongest commonality with Wafer.
It's not a coding agent running on your device optimizing the kernels, we have a system for writing kernels that can be tuned on the target device automatically. So we write the efficient high level kernel structure with tunable parameters, then it fits to whatever hardware it's actually running on.
taylorhou 13 hours ago [-]
Running an inference network across 11 Macs (Teale.com), so I pointed my orchestrator at the repo.
The autotuner is the real thing - kernel-level search, config budget split by measured time share, winners cached per device/toolchain. Rotating resident weights across layers to dodge the hot-cache trap is a nice touch.
One question on model ranking: the fit scores look like predictions built from cost constants measured on a single M4 Max, not per-device measurements. How do you rank models across genuinely mixed hardware? Does the estimator improve from actual runs over time? That gap between predicted and measured fit is what eats mixed machine fleets alive.
Shameless plug since magnitude's goal is highly relevant to what i'm working on: teale.com - distributed inference across fleets of macs. If you're running local models on more than one box, check it out with your agent!
anerli 6 hours ago [-]
Hi, the model ranking scores are based on whatever machine Magnitude is actually running on. So it will account for your specific memory capacity and performance characteristics to recommend appropriate models. The estimator is purely to help filter and recommend a model. The tuning is the only part that actually effects real performance, and is done automatically whenever you download a new model.
p-e-w 20 hours ago [-]
What is the business model?
anerli 20 hours ago [-]
We envision a future where workloads are hybrid. Average consumer hardware will be able to handle a lot with local models, but you’ll still want to use cloud models for harder tasks. Magnitude will make it seamless to switch between the two, even for the same tasks (without breaking your prefix cache). We’ll charge per token for our inference cloud, using the same efficiencies we unlock for local inference to pass the savings on to you.
nullbio 4 hours ago [-]
Why is everyone talking about Mac like it's the only hardware people use?
3 hours ago [-]
pxtail 3 hours ago [-]
Seems like other vendors are removing themselves from the market by not providing what customers want. It's bad and I don't like it as well as I don't like Apples restricted and walled ecosystem and everything but it is what it is.
andrethegiant 15 hours ago [-]
Congrats on the launch!
Zetaphor 14 hours ago [-]
Please consider adding support for Qwen 3.8 Flash Next
anerli 14 hours ago [-]
Will be adding this one very soon!
yolandac 20 hours ago [-]
does it allow us to run larger models that weren't possible before?
anerli 19 hours ago [-]
Right now, since we use less memory for KV, you have more room for model weights when you're running longer sessions.
However we also have expert streaming on the roadmap. This will let you run mixture-of-experts models with unused experts offloaded to RAM or disk, and load them only when needed. This means you'll be able to run models that wouldn't otherwise fit in your GPU memory.
Then it rebuilds latest llama.cpp, grabs the PRs it finds relevant to test against, and then it performs a benchmark and finalizes the upgrade and verifies what model, variant, or even a separate finetune that we should be running.
Occasionally, it performs its own optimizations and commits, which then gets superseded by pull requests and merged code that essentially validates the model's own optimization directionality.
Definitely still helps to reference relevant academic work as well, or even just encouraging the agent to make bigger structural leaps, otherwise it will often get stuck working on low impact micro-optimizations.
I found it to be a good baseline, but at least on Mac there was always something way faster, and/or with better memory requirements - like you said, ds4, omlx, mtplx, etc. It seems if you use local LLMs for real, there is very little reason not to use one of the more optimized engines.
3 main failure modes I observed in the engines:
* Not using best available spec decoding
* Using too much VRAM for KV cache (e.g. KV cache used to take almost nothing in ds4, but huge amount of VRAM on unsloth/llama.cpp for deepseek models)
* Degraded performance at large context sizes - benchmarks at 4K or 32K are awesome, but at realistic 100-200K it's slower than some stupid baseline
Spec decoding: Models in our catalog come assigned with an assigned drafter model for speculative decoding based on the best known method and model available for that target model (support DFlash, DSpark, and DFlash2).
Using too much memory for KV cache: We use a TurboQuant-inspired quantization of KV cache to 8-bit keys and 4-bit values. This drops KV memory usage by over half and also speeds up decode. Based on long context quality benchmarking we've done it does not seem to negatively impact retrieval or coherence over long context.
Large context sizes: our KV quantization helps a lot for this, and we focus our optimizations on specifically longer-context requests since that's what most agent inference actually looks like.
Quantization is another thing. There are so many engines launched with claims about speed, but in many cases it's optimizing specific lower-quality quants. When you have enough resources, you usually want something like W8A16 + full precision KV cache working as fast as possible, not yet another W8A8 or W4A16.
In general, it seems new models are released so fast now - engines don't always have time to really polish the implementation before the next model is released
All our benchmarks are open source so you can check it out here if you'd like: https://github.com/magnitudedev/magnitude/blob/main/inferenc...
Qwen3-4B-Instruct-2507-4bit Qwen3.5-35B-A3B-4bit Qwen3.5-9B-MLX-4bit Qwen3-Reranker-0.6B-4bit Qwen3-Coder-30B-A3B-Instruct-4bit qwen2.5:0.5b
and the results are what i kinda expected to begin with, this adds next to nothing? - also the repo was pivoted from a playwright sub assembly to this not long ago - so my conclusion - THIS MIGHT be worth some watching if you have a model where no-one!, has optimized it at all - and where it does not use anything native to your platform.
results (averages)
VIA rapid-mlx :8902 (MLX) Decode: 175 tok/s TTFT: 64 ms prefill (~760 tok cold): 594 ms
Magnitude 0.2.1 (GGUF/llama.cpp+Rust) Decode: 161 tok/s TTFT: 111 ms prefill (~760 tok cold): 669 ms
https://github.com/raullenchai/Rapid-MLX
Second, more important for agents: decode speed is rarely what hurts. It's resending the same system prompt plus tool schemas every turn. Does self-optimizing cover prefix cache reuse across requests, or is it kernel and layout tuning only?
And is the endpoint OpenAI-compatible? My runtime already talks to llama.cpp and LM Studio through that wrapper, so drop-in is the difference between trying it tonight and not.
Prefix cache is re-used with a prefix tree structure for maximal re-use across sessions sharing prompts.
The endpoint is standard OpenAI compatible chat completions.
Is it a lack of optimizations, or incorrect numbers, or a benchmark artifact (e.g. something which is harder for spec decoding than usual agentic sessions)?
For m5 - there may be some issue with the Metal 4 matmul hardware utilization that could be causing this to be behind here. Will look into this.
Related to the post above: Similar results here M5 Max running Qwen3.8 UD-Q6-K-XL with zlab’s Dflash2 as the drafter.
Both prefill and decode are roughly 2x faster when served from llama.cpp (b10853 or newer) than magnitude 0.2.1.
llama.cpp had that same matmul gap (pre-fill operations) for the M5/A19 and newer silicon until that sha mentioned above.
External/policy-based throttling for temperature control. Unthrottled, my laptop bottom goes skin-burn hot. But fixed compute caps can have non-linearly dreadful performance impacts in particular cases. Plan is a runtime knob, to replace manual limits-kludgery.
I'll use models which barely fit in VRAM+RAM, and are order-1 tok/s slow. So tool call step overhead can be painful - a world where `ls` costs tens of seconds. Plan is blending harness plugins with inference loop, for "no, don't stop - I already have the call result for you - just keep going" (and also some logit games).
https://github.com/magnitudedev/magnitude/issues
Unfortunately even with my 5070ti, llama.cpp seems to be about 20-30% faster at decode, running as:
set CUDA_VISIBLE_DEVICES=0 build\bin\Release\llama-server -hf google/gemma-4-12B-it-qat-q4_0-gguf -ngl 99 --no-mmproj-offload -mg 0 -c 262144 -fa on --host 0.0.0.0
Currently we don't support multi-GPU setups, that is on our near-term roadmap. It saying the model is too big for that GPU might be a bug - would you be willing to open a github issue with more detail on your setup? https://github.com/magnitudedev/magnitude/issues
As for performance, there may be some variability still depending on the model and backend. We have room for improvement for various setups that we are closing as we work out some details with our kernels and tuning system, so appreciate the data point and will look into that combination.
would love to try it again when you have updates
Q: From my (very, very limited!) understanding, I’m under the impression that part of the “inference engine inertia” is that model- or at least architecture-specific code is required for most, if not each new open-weight model coming out.
Assuming I got that right, do you plan on supporting everything vLLM/llama.cpp can do, such that Magnitude becomes a drop-in replacement for as many (economically/pareto-viable) models as possible, or do you want to focus on the best possible support for only a select few models/classes of models?
We plan to support any model architecture that we believe is somewhere along or close to the pareto frontier. There's some model families that are outdated or more niche that we don't necessarily want to put our focus into.
llama.cpp does not saturate memory bandwidth for single-stream tok/s, and for long context and batching, our quantized KV and associated decode kernels allow us to reduce the effective bandwidth needed, and surpass llama.cpp significantly in decode speeds.
For llama.cpp, we try to make the comparison as fair as possible by using similar settings. No speculative decoding, default prefill batch sizes, flash attention on.
We tried also quantizing the KV cache to 8-bit keys and 4-bit values like we do in Magnitude, but this bombed decode speed for llama.cpp in our testing. Since it seems llama.cpp did not optimize that path, we used 16-bit KV instead.
The source for the benchmark is available here also: https://github.com/magnitudedev/magnitude/tree/main/inferenc...
--
[0] https://developer.apple.com/videos/play/wwdc2026/326/
The hard part isn't the optimization itself—it's measuring whether a change actually helps across the long tail of request patterns. Most inference benchmarks show huge gains on popular workloads, but production traffic has a fat tail where naive optimizations hurt latency.
Interested in how you handle regression detection when the optimizer changes between requests.
And a more general question: does your engine detect and optimize for custom setups like multiple (possibly different) GPUs, eGPUs,...? Because if all you have is a stock major system like a Mac or DGX Spark, that's all you're going to care about, and there are a lot of highly optimized single-hardware engines out there that will be hard to beat in the long run. Something that automatically adapts to custom systems that don't have their own subreddits could really fill a gap.
Qwen3.8-Flash-Next support will also be added very soon.
Taking full advantage of all the hardware on your machine in the most performant way possible is the overall goal of the inference engine. This includes a lot of what you're describing. We want to map out the full hardware topology of your system (one or more GPUs, CPU, memory), and compile a combination of kernels to serve a given model optimally across that stack, allocating different parts of the workload wherever it fits best.
Currently we're writing tunable kernels that optimize themselves for one device, but we're working on a kernel compiler that will be able to compile and distribute kernels across any number of devices in a system.
So it's more about making those GPU kernels perform the required memory move and arithmetic operations as close to the theoretical optimum as possible.
Magnitude is optimized for maximum single-session performance and memory efficiency - so we should be more performant for local inference use cases.
Regarding model variants - our catalog includes different quantizations, and automatically assesses these against your hardware to determine which ones will fit in your memory and how fast they will run. This lets you pick a model to download based on your desired speed/intelligence tradeoff.
Maybe have it run silently in the background and assess on demand when a user selects / attempts to download a model. It's not quite clear why all need to be assessed before I can download the first model to try.
Could you share your hardware and OS details to help us identify what might be the issue here?
There's also a github issue open on this topic if you want to leave a comment there: https://github.com/magnitudedev/magnitude/issues/142
Can it do all the shenanigans that allows to run qwen flash on 12GB vram over 40 toks like people seems to be getting in this thread?: https://www.reddit.com/r/LocalLLaMA/comments/1wp7zyb/qwen38f...
Our plan to enable running bigger models on less GPU memory in a way that'll remain productive is expert streaming. This will let you offload experts for MoE models to RAM or disk, and load them when needed. This can have some performance tradeoff, but is lossless.
I am building propelcompute.com an open router for private hardware and experimented with exo labs to run large models on for Mac studios and plan on doing the same with nvidia and amd. Id love to integrate your inference engine into the system but built gpu is critical.
On a 64gb Ram and 5080 machine the biggest model suggested was Qwen 9B
can hit 90+ tps on the MoE 35b but mag thinks it wont fit.
The MoE 35b might be tight though unless you were to go below 4-bit. Could you share the quant you used when you ran this on that 5080 before? Our catalog only contains models down to 4-bit because we find that thinking, tool calling, and overall capabilities start to suffer at lower fidelity.
Feel free also to create a GitHub issue with more details and we can take a closer look.
https://github.com/magnitudedev/magnitude/issues
Let me know if you keep running into problems for some reason
It's not a coding agent running on your device optimizing the kernels, we have a system for writing kernels that can be tuned on the target device automatically. So we write the efficient high level kernel structure with tunable parameters, then it fits to whatever hardware it's actually running on.
The autotuner is the real thing - kernel-level search, config budget split by measured time share, winners cached per device/toolchain. Rotating resident weights across layers to dodge the hot-cache trap is a nice touch.
One question on model ranking: the fit scores look like predictions built from cost constants measured on a single M4 Max, not per-device measurements. How do you rank models across genuinely mixed hardware? Does the estimator improve from actual runs over time? That gap between predicted and measured fit is what eats mixed machine fleets alive.
Shameless plug since magnitude's goal is highly relevant to what i'm working on: teale.com - distributed inference across fleets of macs. If you're running local models on more than one box, check it out with your agent!
However we also have expert streaming on the roadmap. This will let you run mixture-of-experts models with unused experts offloaded to RAM or disk, and load them only when needed. This means you'll be able to run models that wouldn't otherwise fit in your GPU memory.