Point it at a v1 endpoint, or use the internal one. Skip the setup scripts, docker stacks, and broken dependencies. Drop in a model. Ask questions about your codebase, your documents, your spreadsheets. It reads them all, remembers the context, talks back to you, and operates exactly where you tell it to. Your hardware, or their cloud. Your models, or theirs. You decide.
We built llama-server into StackRAG. It manages your local hardware, handles auto-population, and lets you hot-swap models on the fly with zero configuration required. For power users, StackRAG remains completely customizable, giving you effortless control over llama-server's parameters whenever you need it.
Most "local AI" setups are a wonky mess of terminals, config files, and half-broken Python scripts. StackRAG is the finished product. Every subsystem below is built, tested, and working together out of the box.
Drop in your entire codebase, your research papers, your SQL schemas, your meeting notes. Then ask: "Where do we handle auth token refresh?" or "Summarize the risk section of the Q3 report." It finds the exact passages, reads them in context, and answers.
Hold the mic button, dictate your question in natural language, and hear the answer read back to you. Real speech-to-text and text-to-speech, running locally. No audio leaves your machine. Hands-free coding sessions, late-night research, accessibility.
Attach a screenshot of a UI bug, a whiteboard photo, a diagram. Draw a box around the region you care about. The model focuses its analysis on that exact area. The image never gets stored in your database — it's a one-shot context overlay.
Every conversation is a session file. Save it, reload it later, or send the JSON to a teammate. They open it, see the full conversation, and can continue from where you left off. The RAG context is stripped on save — only the conversation travels.
20+ built-in personas: Forensic Coder, Contract Auditor, HVAC Sizing Engineer, SQL Optimizer, Meeting Minutes Writer, Tax Strategist. Each has its own system prompt, temperature, and tone. Or write your own. Editable in-app, no code required.
Bundle a GGUF model, bind a vision projector, and you have a full inference server in under 60 seconds. CUDA or ROCm. Flash Attention. Speculative decoding. 8-bit KV cache. All toggled from one settings panel. No Docker. No Kubernetes. No "it works on my machine."
This isn't a wrapper around an API. Every layer below was designed with its failure mode in mind. The sort order protects the KV-cache. The atomic writes protect the server. The thread limits protect your thermals. The firewall protects your index. Read it if you care about the details.
PDFs, DOCX, PPTX, notebooks, CSVs, source code, configs. Digital PDFs extract locally via PDFium. Scanned PDFs route to a Docling OCR microservice. Jupyter cells are labeled. Tables are flattened to text. Code is split.
Files become chunked parents (context for the LLM) and children (targets for vector search). Code uses language-aware splitters that respect function and class boundaries. This prevents context loss at seams.
Nomic Embed Text runs as an ONNX model on your CPU. Native, mask-aware mean pooling, finalized with strict L₂ normalization. Advanced arena-allocation memory management cuts out runtime allocation bottlenecks.
Vector search hits children (precise). Parent IDs are extracted. Parents are fetched (context). Results are sorted deterministically. This sort order is what makes prefix caching work.
The retrieval results are sorted deterministically. This guarantees the context string is byte-identical across repeated queries against the same files. llama.cpp's automatic prefix caching then is more likely to skip recomputing the K/V tensors for the entire vault context. Follow-up questions on the same files respond in a fraction of the time. The UI shows you a live "PREFIX_STABLE" indicator so you know the cache is hot.
Python, JavaScript, TypeScript, Go, Rust, Java, C/C++ and other supported language files are split using character aware splitting. The splitter parses the language grammar protecting your functions. Docstrings and type hints travel with their function.
No thread is ever force-killed during normal operation. The HALT button sets a flag. The worker checks the flag at every retry boundary, every stream chunk, and every audio callback. If the flag is set, the loop exits cleanly. Only if a thread is trapped in a C-level blocking call does the system escalate to terminate.
Image bytes are intercepted at two enforcement points: the file-selection router and the vault ingestion entry. They are stored only in a transient controller variable, deep-copied into the worker payload, merged into the message as an OpenAI content-parts array, and wiped the moment the SSE stream dispatches. The LanceDB index, the ONNX embedding pipeline, the session JSON, and the permanent message history remain text-only. An image can never contaminate your vector search.
Every save operation is completely all-or-nothing, ensuring your data is never left in a corrupted or half-written state. The server automatically checks and verifies that all required components are fully intact before making any changes. If any external file is missing or deleted while you are working, the application safely halts the operation to prevent data loss or corruption.
If a network interruption occurs, the system automatically performs a full connection reset to clear temporary errors and ensure a clean re-establishment. This process prevents persistent connection errors caused by lingering sessions after your network recovers.
No Docker. No conda. No "install these 47 dependencies in the right order." You have a GPU. You have a GGUF file. That's all you need.
The app launches. Splash screen. License check. Whisper and Kokoro load in the background. UI is ready. Your vault directory is created at ~/Documents/StackRAG/.
Click the Models Manager button. Drop your .gguf file in. If you have a vision projector (*mmproj*.gguf), drop that in too. Select both. Click "Save & Bind." Done. After binding, the vision model will be available to use.
Select AUTO (it detects CUDA or ROCm), set your context size, hit START. The engine handles the rest, spawning a bundled llama-server binary with a locked-in, high-performance configuration: Flash Attention, speculative decoding, 8-bit KV cache, and total GPU layer offloading. No tweaking, no breaking—just raw local speed right out of the box.
The model list fetches from /v1/models. The connection fields auto-update to 127.0.0.1:8080. The config persists to disk. You're talking to your own GPU. The status indicator glows green. Type a question. Get an answer.
Built in push-to-talk speech-to-text via faster-whisper. Streaming text-to-speech via Kokoro. The AI starts talking while it's still generating. Emojis and markdown are stripped before TTS. Interruptible at any time. No audio ever leaves your machine.
Hold the recording button, speak, and release. The app automatically processes your audio input in real time, optimizing clarity and removing silence. Your spoken words are instantly converted into text and added to your input box—or sent automatically if you are in Chat mode.
Your question hits the RAG pipeline. Vault context is retrieved. The model reasons (you see the thought log stream in amber). The answer begins generating. Sentence by sentence, it's queued for speech.
Kokoro speaks the first sentence shortly after generation starts. You hear the answer while the rest streams in visually. Click HALT to stop instantly. The waveform strip in the sidebar shows live audio levels.
The Settings window is a full configuration surface. You don't edit JSON files. You don't hunt through 14 separate config scripts. Everything is in one place, grouped into sections, with live validation and instant apply.
Host, port, base path. Local, LAN, or cloud. The app auto-detects whether to send auth headers. "Connect" button fetches the live model list from the endpoint.
GPU mode (Auto / CUDA / ROCM), context size, START / STOP buttons, and the Models Manager. The server state persists across dialog close/reopen. Stop guarantees GPU VRAM is fully released before the next launch.
Chat history folder, vault directory. All under ~/Documents/StackRAG/ by default. Change them if you want. The app creates the directories on save.
Max tokens for Chat mode and Deep Think mode. The token counter in the UI uses these as its denominator. Overflow turns the counter red. The model's actual context window is set via the llama-server context box.
Background color, panel color. Applied live to the entire application. The cyberpunk terminal theme is the default, but you can tone it down or shift the palette.
Docling OCR service URL (for scanned PDFs). Toggle for logprobs in API responses. Toggle for custom temperature override. Both are "not supported by all endpoints" — the app handles the fallback gracefully.
Stored in the OS credential manager (Windows Credential Manager, macOS Keychain, Linux Secret Service). Never written to disk in plaintext. Never in config.json. Never logged. Required only if you point the app at a cloud endpoint.
Apply changes without closing the app. The config module is hot-reloaded. The API manager re-initializes. The embedding service re-warms. No full restart needed for most changes.
Switch the persona dropdown and the system prompt, temperature, and input placeholder all change. The same model, the same vault, a completely different expert in your corner. Or open the Settings → Personas tab and edit an existing one. No code. No restart.
| DIY Terminal Stack | Ollama Web UI | StackRAG | |
|---|---|---|---|
| Local RAG with your own files | Build it yourself | No | ✓ Parent-child LanceDB |
| Character-aware code splitting | Manual LangChain config | No | ✓ Language-specific |
| Local embeddings (no cloud) | Wire up yourself | No | ✓ ONNX, thread-limited |
| Prompt prefix caching (fast follow-ups) | Unlikely to set up | No | ✓ Deterministic sort + hash |
| Voice input (STT) | Separate tool | No | ✓ Push-to-talk Whisper |
| Voice output (TTS, streaming) | Separate tool | No | ✓ Kokoro, interruptible |
| Image annotation (circle & comment) | Custom build | No | ✓ 0-1000 grid, VL-compliant |
| Session save / load / share | Manual JSON | No | ✓ Forensic RAG scrub |
| Custom personas (editable in-app) | Edit prompt files | No | ✓ 20+ built-in, editable |
| One-click llama-server with vision | Compile + config | No | ✓ Bundled binary, < 60s |
| No cloud dependency for core ops | ✓ (if you set it up) | ✓ | ✓ |
| Single executable, no dependencies | Python + 47 packages | ✓ | ✓ Nuitka standalone |
One file. Double-click. Point at your GPU. Drop in your files. Ask it things. No subscription. No data leaving your machine. No "contact our sales team."
Coming soon — v2.1Last updated: 10-1-26 · Version 1.0 · Governing jurisdiction: USA
By downloading, installing, or using StackRAG (the "Software"), you ("User," "you," or "your") agree to be bound by these Terms of Service, the End User License Agreement (EULA), and the AI Output Disclaimer (collectively, the "Agreement"). If you do not agree, do not install or use the Software. You must be at least 18 years of age or the age of majority in your jurisdiction.
StackRAG is a desktop application for local Retrieval-Augmented Generation (RAG), voice interaction, multimodal image annotation, and session management. It is sold as a single-user, single-machine license. The Software communicates with a local LLM inference server (llama-server) and a local vector database (LanceDB). It does not require a cloud account, does not transmit your files to a third party, and does not require a subscription. All data processing occurs on the User's own hardware unless the user specificially points StackRAG to a cloud subscription endpoint.
Subject to payment and compliance with these Terms, Stackraglabs llc ("we," "us," "our") grants you a non-exclusive, non-transferable, non-sublicensable, revocable license to install and use one (1) copy of the Software on a single computer that you own or control. This license is personal to you. It does not extend to team members, clients, contractors, or any third party. See the EULA for full terms.
You agree NOT to:
The Software, including its code, UI design, icons, documentation, and all associated materials, is the proprietary property of Stackraglabs llc and is protected by copyright, trade secret, and other intellectual property laws. The license key and machine-binding mechanism are proprietary. You do not acquire any ownership interest in the Software by purchasing a license. All rights not expressly granted are reserved.
The files you index, the conversations you have, the audio you record, and the images you annotate are your data. They are stored locally on your machine. We do not collect, transmit, store, or have access to your data. You are solely responsible for the content you feed into the Software and for any decisions you make based on its output. See the Privacy Policy for details on the minimal operational data that is processed (license activation, OS credential storage).
The Software incorporates and bundles open-source components, including but not limited to: llama.cpp (MIT), PySide6 / Qt (LGPL v3), LanceDB (Apache 2.0), faster-whisper (MIT), Kokoro TTS (Apache 2.0), ONNX Runtime (MIT), and HuggingFace Transformers (Apache 2.0). Your use of these components is governed by their respective licenses, which are summarized in the Open Source Components section of this site. The inclusion of open-source components does not grant you any rights in Stackraglabs llc proprietary code, UI, or architecture.
We may provide updates, patches, and bug fixes at our discretion. We do not guarantee a specific update schedule, feature roadmap, or support SLA. Support is provided on a best-effort basis via email at [email protected]. We are not obligated to provide support for: (a) hardware that does not meet the stated system requirements; (b) issues caused by user modification of the Software or its configuration; (c) issues arising from use in an unsupported OS or GPU configuration; (d) issues caused by third-party models or endpoints that the User configures externally.
We may suspend or terminate your license immediately, without notice, if you breach any term of this Agreement. Upon termination, you must cease all use of the Software and destroy all copies. Sections 6, 11, 12, 13, and 14 survive termination.
THE SOFTWARE IS PROVIDED "AS IS" AND "AS AVAILABLE," WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE, NON-INFRINGEMENT, ACCURACY, COMPLETENESS, OR ERROR-FREE OPERATION. WE DO NOT WARRANT THAT THE SOFTWARE WILL BE UNINTERRUPTED, SECURE, OR FREE OF DEFECTS. WE DO NOT WARRANT THAT AI-GENERATED OUTPUT WILL BE ACCURATE, COMPLETE, CURRENT, OR APPLICABLE TO YOUR SPECIFIC CIRCUMSTANCES.
TO THE MAXIMUM EXTENT PERMITTED BY APPLICABLE LAW, IN NO EVENT SHALL Stackraglabs llc, ITS DEVELOPERS, AFFILIATES, OR LICENSORS BE LIABLE FOR ANY INDIRECT, INCIDENTAL, SPECIAL, CONSEQUENTIAL, OR PUNITIVE DAMAGES, OR ANY LOSS OF PROFITS, REVENUE, DATA, OR GOODWILL, ARISING OUT OF OR IN CONNECTION WITH THE USE OF OR INABILITY TO USE THE SOFTWARE, REGARDLESS OF THE THEORY OF LIABILITY (CONTRACT, TORT, STRICT LIABILITY, OR OTHERWISE), EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGES. OUR ENTIRE AGGREGATE LIABILITY FOR ALL CLAIMS RELATED TO THE SOFTWARE SHALL NOT EXCEED THE AMOUNT YOU PAID FOR THE LICENSE (ONE-TIME PURCHASE PRICE). SOME JURISDICTIONS DO NOT ALLOW THE EXCLUSION OF CERTAIN WARRANTIES OR LIMITATIONS OF LIABILITY, SO SOME PROVISIONS MAY NOT APPLY TO YOU.
You agree to indemnify, defend, and hold harmless Stackraglabs llc, its officers, employees, and agents from and against any and all claims, damages, losses, liabilities, and expenses (including reasonable attorneys' fees) arising out of or related to: (a) your use or misuse of the Software; (b) your violation of these Terms; (c) any content you generate or process using the Software; (d) any decision you make based on AI output; (e) any breach of third-party rights by your use of the Software.
These Terms are governed by the laws of the USA, without regard to its conflict-of-law provisions. Any dispute arising under this Agreement shall be resolved exclusively in the courts of Colorado county TX. You consent to personal jurisdiction and venue therein. If any provision of these Terms is held to be invalid or unenforceable, the remaining provisions shall continue in full force and effect.
We may update these Terms from time to time. The "Last updated" date above reflects the current version. By continuing to use the Software after changes are posted, you agree to be bound by the revised Terms. If we make material changes, we will notify you via email prior to the effective date.
Questions about these Terms? Contact: [email protected]
Stackraglabs llc
StackRAG v2.x · Single-User Desktop License · [DATE]
This license grants you the right to install and run StackRAG on one (1) physical computer that you own or exclusively control. The license is bound to your machine. The license is non-transferable. You may not install the Software on a second machine, a virtual machine, a cloud instance, or a shared workstation without purchasing an additional license.
After initial online activation your license does not expire, does not require renewal, and does not depend on network availability.
The Software bundles pre-compiled binaries of llama-server (MIT License) for local LLM inference. These binaries are included for your convenience and are subject to the MIT License. You may not redistribute the bundled binaries separately. The Software also downloads an embedding model (nomic-embed-text-v1.5) from HuggingFace Hub on first use. This download is governed by the model's license (MIT). You are responsible for complying with the terms of any model files you separately download and place in the models directory.
The Software is compatible with GGUF model files and OpenAI-compatible v1 API endpoints. However, we do not provide support, warranty, or liability coverage for: (a) models you download from third parties; (b) endpoints you configure externally (cloud APIs, LAN servers); (c) output quality or accuracy of any specific model; (d) licensing terms of third-party models (e.g., a model with a non-commercial license used in a commercial context is your responsibility).
Session files (JSON) saved by the Software contain conversation history with RAG context stripped. You are free to share session files with others for review, collaboration, or reference. However, sharing a session file does not constitute a license grant for the Software itself. The recipient must purchase their own license to use StackRAG.
Upon termination (by you or by us), the license key becomes invalid, the Software will no longer pass its license check, and you must uninstall the Software and delete all copies, including the local vault directory, chat history, and license file. Sections 5, 6, and the Limitation of Liability / Indemnification clauses survive termination.
Last updated: 10-1-26 · Version 1.0
Your files, your conversations, your audio, your images — they stay on your computer unless you choose to use a cloud endpoint. We do not have a server that receives your data. We do not run analytics on your usage. We do not sell, share, or aggregate your information.
The following third parties may process data in connection with the Software. We are not responsible for their data handling:
Since we do not collect or store your operational data, there is no retention policy to enforce. The local files on your machine (vault, chat history, license file, token cache) are under your control. You can delete them at any time by removing the ~/Documents/StackRAG/ directory and the license file. The activation server retains the license activation record (key, instance ID, timestamp) for the duration of the license.
If you are located in the European Economic Area, United Kingdom, or California, you have the right to:
To exercise any of these rights, email [email protected].
This website DOES NOT use cookies. This website does not use cookies, tracking pixels, or analytics scripts. No browsing data is collected.
The Software is not directed at individuals under the age of 13 (or 16 in the EEA). We do not knowingly collect data from children.
We may update this Privacy Policy. Material changes will be announced via email to registered license holders. The "Last updated" date reflects the current version.
Privacy questions: [email protected]
Last updated: [DATE]
StackRAG is a digital product delivered instantly upon purchase. Once the download is available and the license key is issued, you have 10 days to return it. we stand behind our products.
Email [email protected] with:
We will respond within 48 hours. The refund is processed through Polar.sh to your original payment method within 5–10 business days. You must uninstall the Software and delete the license file upon refund.
If you dispute a charge with your card issuer or bank (a "chargeback") rather than requesting a refund through us, we will respond to the dispute through Polar.sh. If the chargeback is upheld, your license is immediately revoked and the Software will no longer function. Repeated or fraudulent chargebacks will result in a permanent ban from purchasing our products.
Read this before using professional-advice personas. 10-1-26
StackRAG is a software tool. It is not a professional.
The Software uses large language models (LLMs) to generate text output. LLMs are probabilistic systems. They can produce confident, well-formatted, and completely wrong answers. They do not "know" the National Electrical Code. They do not "know" ICD-10 billing rules. They do not "know" your jurisdiction's tax code. They pattern-match on training data and generate plausible text. Plausible is not the same as correct.
The Software includes personas that frame output in the language of specific professions:
You are solely responsible for every decision you make based on the Software's output. You are responsible for verifying any factual claim, code reference, calculation, or recommendation against primary sources (code books, tax statutes, medical guidelines, engineering handbooks, your jurisdiction's regulations) before acting on it. You are responsible for consulting a licensed professional in the relevant field before making any decision that carries legal, financial, safety, or health consequences.
Stackraglabs llc makes no representation or warranty, express or implied, that any output generated by the Software is accurate, complete, current, applicable to your specific situation, or free from error. LLMs can "hallucinate" — generating text that is confident and well-structured but factually incorrect, citing non-existent statutes, misquoting code sections, or applying outdated regulations. This is a known and inherent limitation of the underlying technology. We do not monitor, verify, or guarantee the accuracy of AI output.
If you use the Software's output in a professional capacity (electrical work, medical billing, tax preparation, legal filing, engineering design, construction) and that output results in a loss, liability, fine, penalty, injury, or damage — to you, your client, your employer, or any third party — you bear that risk. Stackraglabs llc is not a licensed professional, does not carry professional liability insurance, and is not a substitute for the professional judgment of a qualified expert in the relevant field. You agree to indemnify Stackraglabs llc against any claims arising from your use of professional-advice output.
By using the Software, you acknowledge that you have read and understood this Disclaimer. You understand that the Software is a productivity and information tool, not a professional service. You assume full responsibility for your use of its output.
Third-party software included with or used by StackRAG.
StackRAG incorporates the following open-source and third-party components. Your use of these components is governed by their respective licenses. Full license texts are available at the linked URLs.
| Component | License | Purpose |
|---|---|---|
| llama.cpp / llama-server | MIT | Local LLM inference engine |
| PySide6 (Qt for Python) | LGPL v3 | GUI framework (all UI widgets) |
| LanceDB | Apache 2.0 | Vector database (local vault) |
| faster-whisper | MIT | Speech-to-text (Whisper CTranslate2) |
| Kokoro TTS | Apache 2.0 | Text-to-speech synthesis |
| ONNX Runtime | MIT | Local embedding inference |
| HuggingFace Transformers / Tokenizers | Apache 2.0 | Tokenization |
| nomic-embed-text-v1.5 | MIT | Embedding model |
| pypdfium2 | Apache 2.0 | PDF text extraction |
| python-docx / python-pptx | MIT | Office document parsing |
| pandas / numpy | BSD-3-Clause | Data processing, CSV/Excel parsing |
| OpenAI Python SDK | MIT | Cloud LLM integration |