LLM speculative inference server for heterogeneous hardware & consumer GPUs
-
Updated
Sep 1, 2026 - C++
LLM speculative inference server for heterogeneous hardware & consumer GPUs
BestBuy Bot is an Add to cart and Auto Checkout Bot. This auto buying bot can search the item repeatedly on the ITEM page using one keyword. Once the desired item is available it can add to cart and checkout very fast. This auto purchasing BestBuy Bot can work on Firefox Browser so it can run in all Operating Systems. It can run for multiple ite…
Pre-built wheels for llama-cpp-python across platforms and CUDA versions
Best Buy Bullet Bot, abbreviated to 3B Bot, is a stock checking bot with auto-checkout created to instantly purchase out-of-stock items on Best Buy once restocked. It was designed for speed with ultra-fast auto-checkout, as well as the ability to utilize all cores of your CPU with multiprocessing for optimal performance.
Run Qwen3.8 Flash Next on dual RTX 3090 (2x24GB) + 128GB RAM: 256K context, 75.6 tok/s at 258K input + 4K output (3 runs), 135 tok/s selected warm short-prompt decode. Pinned vLLM, CPU offload, MTP, P2P guide. Dual RTX 4090 testing notes (unvalidated).
Open source DIY AI computing platform: Build a powerful RTX 3090 GPU rig under €1,300 for local LLM inference, training, and AI development. Complete with parts list, assembly guide, Ubuntu setup, and remote access configuration. Ideal for students, researchers, and hobbyists wanting affordable AI hardware without cloud dependencies.
🌐 Automate and manage BitBrowser tasks for efficient Google One student discount processing in one streamlined system.
Full benchmark traces: Laguna S 2.1 INT4 vs DFlash speculative decoding on 4x RTX 3090 (vLLM 0.25, TP4, 96GB VRAM). 152 measurements, raw JSON + GPU telemetry.
PlantCare AI | High-performance Plant Disease Detection system optimized for RTX 3090. Identifies 38 diseases with 93% accuracy. Features a futuristic Glassmorphism UI, real-time diagnosis, and AI-driven treatment recommendations. 🌿🌱
poolside Laguna-S-2.1 INT4 + DFlash speculative decoding on 4x RTX 3090: 200K context, 282 tok/s peak decode, gate-proven with a 190K-token prompt. Full levers menu + failure catalog.
Qwen3.8 Flash Next on 4x RTX 3090: full-GPU PLE, CUDA Graph, IQ4_NL, 262K context, multimodal / 四卡3090部署优化
Dynamic open-source Google Sheets tables inspired by the famous Hive Systems ones. Just input the H/s/GPU data.
Benchmark llama.cpp prefill/decode speed at real context depth, not on an empty context. Includes a Qwen3.6-27B run on 1x and 2x RTX 3090.
To associate your repository with the rtx3090 topic, visit your repo's landing page and select "manage topics."