Feat/llamacpp cloud providers - #365
Merged
Merged
Conversation
Replace the Ollama-backed agent with a managed llama.cpp pool (one llama-server per loaded model, LRU + VRAM-budget eviction, idle reaper), a shared model library with resumable downloads, and external providers (OpenAI, Anthropic, Mistral, Groq, OpenRouter, Ollama, custom endpoint) with API keys encrypted through Electron safeStorage. Review fixes: - restart a crashed slot on a fresh port instead of killing the server that inherited its port - unload_all() spares slots answering a request or being started; the agent releases its slot before the post-workflow unload - "Free memory" also unloads the local LLMs - POSIX stale-server cleanup only kills llama-server processes - retry an external chat without images when a text-only model rejects them - only treat OSCrypt-tagged hex as ciphertext so legacy hex keys survive - external model picker no longer shows a model the draft does not hold - drop the CAD and Custom tabs from the model library
scripts/react-test-env.mjs imports jsdom, which no package declared, so llmModelsStore.react.test.mjs failed on any fresh install and npm test (and with it the pre-push hook) failed for everyone.
…lder The llama.cpp engine, GGUF models, logs and pool config lived in ~/.modly/llm on the system drive, apart from every other data folder the user picked at setup. They now go to an `agent` folder beside models/, extensions/, workflows/, workspace/ and dependencies/, passed to the API as MODLY_LLM_DIR. Setup writes <base>/agent. Installs that predate it get the folder next to their models folder, created and pinned in settings.json at startup so moving the models folder later does not take the agent folder with it.
Settings > Agent shows whether the llama.cpp engine is installed, with an install button when it is missing, and lists the models in the agent folder with the selected one first. Browse opens the model library, now a wide dialog with two sections: Installed (everything in agent/models, Select / delete) and Suggested (catalog models not downloaded yet). Add picks a local .gguf and copies it into the agent models folder; picking and copying both run in the main process. The engine install moved out of the dialog into the settings section, and the shared SSE progress bar lives in its own component.
Agent settings now use the Settings card kit in two columns: provider, local engine (status badge, install, selected model, simultaneous models) and thinking on the left; the MCP server setup on the right, with copy buttons and a tab per client. The shared Card gains an optional `aside` slot for header badges. The model library shows models as cards with a search box and filters (All, Fits my GPU, Vision, CAD), the card VRAM next to Add, and per-card tags, VRAM estimate and actions (Select, Download, Pause/Resume, Cancel, delete).
The llama-server archive is downloaded from the llama.cpp GitHub release and its files are executed, but nothing checked the bytes. GitHub reports a sha256 digest for every release asset; the download is now hashed as it streams and refused on a mismatch. An asset without a digest (older uploads) is accepted as before. A failed or refused download no longer leaves its temp archive behind.
The chat leaves code/CAD models out of its own picker, but the model library still offered Select on them, so a CAD coder could become the agent's chat model. Those cards now say they are for workflow nodes. Also drop LlmModelSelect, which nothing imports.
lightningpixel
approved these changes
Oct 3, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.