Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
localllm
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
VRAM and RAM for local LLMs — honest planning bands, not a GPU tier list
Tech-Gurunomics
Tech-Gurunomics
Tech-Gurunomics
Follow
Sep 6
VRAM and RAM for local LLMs — honest planning bands, not a GPU tier list
#
localllm
#
llm
#
ollama
#
hardware
Comments
Add Comment
4 min read
Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Sep 5
Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.
#
llm
#
benchmarks
#
localllm
#
reproducibility
1
 reaction
Comments
1
 comment
3 min read
Our 4B beat Claude Opus on a 440K-token corpus. Then it came last on the public benchmark.
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Sep 5
Our 4B beat Claude Opus on a 440K-token corpus. Then it came last on the public benchmark.
#
llm
#
benchmarks
#
localllm
#
privateai
2
 reactions
Comments
Add Comment
4 min read
I told the model to separate fields with <TAB>. It did exactly that, and I lost 79 percent of my data.
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Sep 5
I told the model to separate fields with <TAB>. It did exactly that, and I lost 79 percent of my data.
#
llm
#
prompting
#
debugging
#
localllm
1
 reaction
Comments
Add Comment
3 min read
A 4B on a 6GB laptop matched frontier-model accuracy on aggregation — except when the answer is a number
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Sep 5
A 4B on a 6GB laptop matched frontier-model accuracy on aggregation — except when the answer is a number
#
llm
#
benchmarks
#
localllm
#
agents
1
 reaction
Comments
Add Comment
4 min read
Your agent truncates the corpus and answers anyway. Two harnesses, and a router that picks between them.
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Sep 5
Your agent truncates the corpus and answers anyway. Two harnesses, and a router that picks between them.
#
llm
#
agents
#
localllm
#
benchmarks
1
 reaction
Comments
Add Comment
5 min read
What really fits in 8GB VRAM
the kilted dev
the kilted dev
the kilted dev
Follow
Aug 18
What really fits in 8GB VRAM
#
localllm
#
vram
#
gpu
#
buildinpublic
Comments
Add Comment
7 min read
Moving Scheduled LLM Curation from Cloud APIs to Local Models
Guatu
Guatu
Guatu
Follow
Aug 14
Moving Scheduled LLM Curation from Cloud APIs to Local Models
#
aiagents
#
localllm
#
ollama
#
kubernetes
Comments
Add Comment
9 min read
Nine ways to talk to a local model
the kilted dev
the kilted dev
the kilted dev
Follow
Aug 11
Nine ways to talk to a local model
#
localllm
#
ollama
#
llamacpp
#
buildinpublic
Comments
Add Comment
9 min read
A 4B model on a 6GB laptop beat Claude Opus on our 440K-token corpus. The fix was giving the model less to do.
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Aug 24
A 4B model on a 6GB laptop beat Claude Opus on our 440K-token corpus. The fix was giving the model less to do.
#
privateai
#
llm
#
inference
#
localllm
1
 reaction
Comments
1
 comment
4 min read
[Day 20] Local AI vs cloud AI: one cat photo, 10 video models
PEPPERCORN
PEPPERCORN
PEPPERCORN
Follow
Aug 6
[Day 20] Local AI vs cloud AI: one cat photo, 10 video models
#
localllm
#
ai
#
dgxspark
#
videogeneration
Comments
Add Comment
3 min read
I A/B tested my own system prompt: 24 generations, one clear win, one rule that did nothing | Flash Onyx 2.2
Nathan C.
Nathan C.
Nathan C.
Follow
Aug 29
I A/B tested my own system prompt: 24 generations, one clear win, one rule that did nothing | Flash Onyx 2.2
#
ai
#
ollama
#
localllm
#
opensource
Comments
Add Comment
8 min read
Flash Onyx 2.2: teaching a local model law and game feel
Nathan C.
Nathan C.
Nathan C.
Follow
Aug 28
Flash Onyx 2.2: teaching a local model law and game feel
#
ai
#
ollama
#
localllm
#
opensource
1
 reaction
Comments
Add Comment
4 min read
Running an LLM agent entirely in your browser
Lajos Bencz
Lajos Bencz
Lajos Bencz
Follow
Jul 15
Running an LLM agent entirely in your browser
#
ai
#
localllm
#
webdev
#
machinelearning
Comments
Add Comment
5 min read
Cloud bills kept climbing from 24/7 AI development — I moved the decisions and the implementation to my own local LLMs and cut the cost
uehara
uehara
uehara
Follow
Aug 15
Cloud bills kept climbing from 24/7 AI development — I moved the decisions and the implementation to my own local LLMs and cut the cost
#
ai
#
llm
#
localllm
#
costoptimization
Comments
Add Comment
10 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account