Skip to content

Feat/llamacpp cloud providers - #365

Merged
lightningpixel merged 8 commits into
devfrom
feat/llamacpp-cloud-providers
Oct 3, 2026
Merged

lightningpixel merged 8 commits into
devfrom
feat/llamacpp-cloud-providers

Conversation

@Lorchie

@Lorchie Lorchie commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

No description provided.

Lorchie and others added 8 commits October 3, 2026 13:43
Replace the Ollama-backed agent with a managed llama.cpp pool (one
llama-server per loaded model, LRU + VRAM-budget eviction, idle reaper),
a shared model library with resumable downloads, and external providers
(OpenAI, Anthropic, Mistral, Groq, OpenRouter, Ollama, custom endpoint)
with API keys encrypted through Electron safeStorage.

Review fixes:
- restart a crashed slot on a fresh port instead of killing the server
  that inherited its port
- unload_all() spares slots answering a request or being started; the
  agent releases its slot before the post-workflow unload
- "Free memory" also unloads the local LLMs
- POSIX stale-server cleanup only kills llama-server processes
- retry an external chat without images when a text-only model rejects them
- only treat OSCrypt-tagged hex as ciphertext so legacy hex keys survive
- external model picker no longer shows a model the draft does not hold
- drop the CAD and Custom tabs from the model library
scripts/react-test-env.mjs imports jsdom, which no package declared, so
llmModelsStore.react.test.mjs failed on any fresh install and npm test
(and with it the pre-push hook) failed for everyone.
…lder

The llama.cpp engine, GGUF models, logs and pool config lived in
~/.modly/llm on the system drive, apart from every other data folder the
user picked at setup. They now go to an `agent` folder beside models/,
extensions/, workflows/, workspace/ and dependencies/, passed to the API
as MODLY_LLM_DIR.

Setup writes <base>/agent. Installs that predate it get the folder next
to their models folder, created and pinned in settings.json at startup so
moving the models folder later does not take the agent folder with it.
Settings > Agent shows whether the llama.cpp engine is installed, with an
install button when it is missing, and lists the models in the agent
folder with the selected one first.

Browse opens the model library, now a wide dialog with two sections:
Installed (everything in agent/models, Select / delete) and Suggested
(catalog models not downloaded yet). Add picks a local .gguf and copies it
into the agent models folder; picking and copying both run in the main
process. The engine install moved out of the dialog into the settings
section, and the shared SSE progress bar lives in its own component.
Agent settings now use the Settings card kit in two columns: provider,
local engine (status badge, install, selected model, simultaneous
models) and thinking on the left; the MCP server setup on the right, with
copy buttons and a tab per client. The shared Card gains an optional
`aside` slot for header badges.

The model library shows models as cards with a search box and filters
(All, Fits my GPU, Vision, CAD), the card VRAM next to Add, and per-card
tags, VRAM estimate and actions (Select, Download, Pause/Resume, Cancel,
delete).
The llama-server archive is downloaded from the llama.cpp GitHub release
and its files are executed, but nothing checked the bytes. GitHub reports
a sha256 digest for every release asset; the download is now hashed as it
streams and refused on a mismatch. An asset without a digest (older
uploads) is accepted as before. A failed or refused download no longer
leaves its temp archive behind.
The chat leaves code/CAD models out of its own picker, but the model
library still offered Select on them, so a CAD coder could become the
agent's chat model. Those cards now say they are for workflow nodes.

Also drop LlmModelSelect, which nothing imports.
@lightningpixel
lightningpixel merged commit 221e44e into dev Oct 3, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants