DEV Community

#localllm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
VRAM and RAM for local LLMs — honest planning bands, not a GPU tier list

VRAM and RAM for local LLMs — honest planning bands, not a GPU tier list

Comments
4 min read
Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.

Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.

1
Comments 1
3 min read
Our 4B beat Claude Opus on a 440K-token corpus. Then it came last on the public benchmark.

Our 4B beat Claude Opus on a 440K-token corpus. Then it came last on the public benchmark.

2
Comments
4 min read
I told the model to separate fields with <TAB>. It did exactly that, and I lost 79 percent of my data.

I told the model to separate fields with <TAB>. It did exactly that, and I lost 79 percent of my data.

1
Comments
3 min read
A 4B on a 6GB laptop matched frontier-model accuracy on aggregation — except when the answer is a number

A 4B on a 6GB laptop matched frontier-model accuracy on aggregation — except when the answer is a number

1
Comments
4 min read
Your agent truncates the corpus and answers anyway. Two harnesses, and a router that picks between them.

Your agent truncates the corpus and answers anyway. Two harnesses, and a router that picks between them.

1
Comments
5 min read
What really fits in 8GB VRAM

What really fits in 8GB VRAM

Comments
7 min read
Moving Scheduled LLM Curation from Cloud APIs to Local Models

Moving Scheduled LLM Curation from Cloud APIs to Local Models

Comments
9 min read
Nine ways to talk to a local model

Nine ways to talk to a local model

Comments
9 min read
A 4B model on a 6GB laptop beat Claude Opus on our 440K-token corpus. The fix was giving the model less to do.

A 4B model on a 6GB laptop beat Claude Opus on our 440K-token corpus. The fix was giving the model less to do.

1
Comments 1
4 min read
[Day 20] Local AI vs cloud AI: one cat photo, 10 video models

[Day 20] Local AI vs cloud AI: one cat photo, 10 video models

Comments
3 min read
I A/B tested my own system prompt: 24 generations, one clear win, one rule that did nothing | Flash Onyx 2.2

I A/B tested my own system prompt: 24 generations, one clear win, one rule that did nothing | Flash Onyx 2.2

Comments
8 min read
Flash Onyx 2.2: teaching a local model law and game feel

Flash Onyx 2.2: teaching a local model law and game feel

1
Comments
4 min read
Running an LLM agent entirely in your browser

Running an LLM agent entirely in your browser

Comments
5 min read
Cloud bills kept climbing from 24/7 AI development — I moved the decisions and the implementation to my own local LLMs and cut the cost

Cloud bills kept climbing from 24/7 AI development — I moved the decisions and the implementation to my own local LLMs and cut the cost

Comments
10 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.