An AI assistant for the admin content editor and site editor, with streaming responses, reasoning display, file attachments, tool calls that can read and edit the page, and an optional server-side proxy that keeps API keys out of the browser.
Settings
Open Plugins > AI Assistant > Settings and save once after enabling. The three fields that matter most are provider, key and model - the rest have working defaults. Bulk-translation overrides (provider, model, key, glossary) live on the separate Translate tab; writer voice, disclosure and stock covers live on the Writer tab.
Provider
Type to search all providers from the models.dev catalog (around two hundred, e.g. OpenAI, Anthropic, Google, DeepSeek, Groq, Mistral, xAI, OpenRouter, OpenCode Zen, NVIDIA NIM). The dropdown option shows the display name and model count.
Unknown slugs stay typable on purpose: a brand-new models.dev provider works for chat immediately, and rebuilding the catalog below adds it to the list with full metadata.
API key
One key for the active provider - the placeholder hint follows the selection (sk-ant-…, AIza…, …). Switching providers keeps each provider's key server-side; only the active one is ever used. Installations that saved per-provider keys before keep working: the old key is carried over the first time you save.
Model
Type to search the selected provider's models. Picking a known model auto-fills context size and max tokens from the catalog and stores a metadata snapshot used at runtime. The info line under the field shows context window, max output, per-1M pricing, reasoning support (including which API dialect a routed model uses, e.g. via responses API), attachment kinds and tool support. Free text stays allowed for ids missing from the catalog - the manual context fields apply then.
Refresh fetches the live model list through the proxy (or directly from the provider in direct mode). Catalog rebuilds the whole local catalog from models.dev on the server (php cli.php admin module=plugins/ai-assistant/settings action=buildModels does the same from the command line) and reloads the current list.
Endpoint override (Url)
Optional custom endpoint for the active provider, e.g. a self-hosted gateway speaking its API. Empty means the catalog endpoint. Changing provider empties the field automatically with a notice, so a stale override can never silently misroute the new provider.
General tab - connection, model, shared behavior
- Api mode -
proxyrelays requests through the site backend so keys never reach the browser;directcalls the provider from the browser (the key travels client-side, and the provider must allow the browser origin - e.g.OLLAMA_ORIGINSfor Ollama). For a local server: direct reaches the LLM on your own computer, proxy reaches the LLM on the web server. - Proxy URL - the address the browser calls for relayed requests. Leave the default unless the admin path is customized; only used in proxy mode.
- Session header - upstream header carrying the chat session id on proxied requests (default
x-vvveb-session), so provider-side logs correlate with a conversation. It never lands in the provider body; empty disables it. - System prompt - appended to every chat request to set assistant behaviour. Keep it short; bulk jobs use their own task prompts plus this line.
- Temperature - sampling randomness for chat (0 = deterministic). Omitted automatically for reasoning models that reject it; bulk translate/write set their own values (0 and 0.7).
- Context size / Max tokens - auto-filled from the catalog when you pick a known model. Context size drives the usage bar and auto-compact; max tokens caps each answer and is clamped to the model's own output limit, so an oversized value fails safe instead of 400ing upstream.
- Image model / provider / key / size - the editor's Generate image from text button and writer cover art. Provider and key empty inherit the main ones above (e.g. chat locally, images via OpenAI); local servers have no images API, so covers need a compatible provider here. Size examples:
1024x1024,1792x1024. - Monthly token cap - max bulk tokens per admin per calendar month (0 = off). Every bulk action checks it before spending and fails fast with used/cap numbers; previews show month-to-date next to the estimate.
- Job webhook - HTTPS URL posted to (JSON, fire-and-forget) when a bulk job finishes. Your own endpoint: automation tools, not the provider.
Translate tab - bulk translation overrides
All four fields are optional and share one rule: empty inherits the main setting (provider → model → key, in that order; task keys resolve server-side so secrets never travel). The classic setup is a cheap fast model here while chat stays on the strong one. Precedence for keys everywhere in the plugin is task key, then active key, then legacy per-provider keys from older versions.
- Translate provider / model / key - used by Bulk Translate, the write → translate pipeline, and translation cron. A model with a catalog entry also brings its output cap and pricing into estimates.
- Translate glossary - one
Original = Replacementper line (e.g.checkout = checkout), always used exactly: bulk translation, pipeline translations, and article writing all inject it. Brand and product names belong here. - Translate reasoning -
Offby default. Sendsreasoning_effort: none/minimalonly when the model's catalog entry declares it and only on bearer chat dialects, retrying plain if a gateway rejects it. Field translation needs no chain of thought, and you stop paying output rates for discarded thinking.
Writer tab - article voice, disclosure, stock covers
- Tone presets - one
Name | instructionper line (e.g.Playful | Write with light humor and vivid examples). The names appear in the writer's Tone dropdown; the text is appended server-side, so clients can only pick, never inject. - AI disclosure text - appended as a footer when Disclose AI is ticked on a run; empty uses the default reviewed-by-editors sentence. Opt-in per run, never silent.
- Stock provider - source order for Stock link covers:
Autotries each configured API key in order with the keyless Unsplash search last; explicit Unsplash/Pexels API or scrape-only are also selectable. - Unsplash / Pexels keys - free tiers (Unsplash developers: 50 requests/hour) enable the official search APIs; empty skips straight to the next source. Every candidate is host-allowlisted, IP-checked, downloaded and image-verified before it touches a post.
- Write reasoning -
Onby default (model behavior unchanged). Turn off once outlines carry your planning - same suppression mechanics as Translate reasoning above.
Using the assistant
- Thinking level (Off/Low/Medium/High) follows the selected model: provider-specific effort values where the catalog defines them, generic levels otherwise, disabled with an explanation when the model cannot reason. Reasoning blocks and tool calls render as collapsible cards with a Running/Done/Failed status pill: thinking streams open and collapses when finished, tool calls stay collapsed throughout (a manual expand pins a card open), errors stay open, and everything hides behind the thinking/tools toggles.
- Attachments - the file picker only offers formats the model takes; anything else is refused with a message naming what is supported. Text files (.txt, .md, .csv, …) are inlined as text so they work on every model including text-only ones; long files are truncated with a marker. Image-capable models also receive page screenshots as vision.
- Costs - every assistant message shows its estimated cost from catalog pricing; the usage footer totals the session and shows cache-hit savings. Encrypted reasoning never renders - it is preserved invisibly so multi-turn tool calls keep working.
- Failures - dropped connections and overloaded providers retry automatically with backoff (1s → 2s → 4s, longer for rate limits) and a live countdown; Stop cancels the wait. Deterministic errors (bad key, unknown model, full context) surface immediately with guidance instead.
- Sessions survive reloads and sync server-side; Compact shrinks a long conversation into a summary when context runs out. When browser storage nears its quota, sessions upload automatically and older local copies trim to metadata (reloaded on open) - a warning replaces the old silent failure.
- The assistant UI follows the admin language through the CMS-wide
Vvveb.i18nruntime (public/js/libs/gettext.js, gettext-compatible plurals): sidebar, manager and editor strings translate from the same catalogs as the backend (translators add the msgids fromplugins/ai-assistant/system/i18n-data.phptolocale/<lang>/LC_MESSAGES/vvveb.po; English always works untranslated). - Article writer (Plugins → AI Article Writer) turns a topic list (one per line, up to 200) into draft posts: pick language, site, categories (multi), tags, length (short/medium/long) and status (draft default - bulk AI content never auto-publishes), preview the call/token estimate, then run the resumable stepper with per-topic results. Uses the active chat model (writing wants the strongest model) with temperature 0.7, JSON mode + repair retry, and core post creation (headline, HTML body, excerpt, meta description/keywords, cover alt, slug, site link). Re-running the same topics skips already-written ones (topic-hash ledger, reported with the original post id). Outlines runs a cheap planning pass first - tick approved topics and Start writes only those. Completion notices report tokens used; the last run summary persists on the page. Cover image switch: No cover, AI-generated (paints one cover per article via the image task settings), or Stock link. Stock mode resolves sources through a provider registry (
system/stock.php- implementStockProviderand add one line to register the next one): the Stock provider setting picks Auto (every configured API key in order, keyless scrape last), Unsplash API, Pexels API, or scrape-only. API results carry photographer credit into the run log; a random hit from the first 10 is verified and stored, failures fall through to a model-suggested URL and then to generation, each step logged. Every candidate is validated (allowlisted CDN host, public IP, real image ≥200×200) before it touches a post; anything suspicious is rejected down the chain. Covers save undermedia/ai-assistant/; a failed cover still saves the post, imageless, with a warning. Stock photos carry the source URL in the log - check the license before publishing. - Drip publishing. Set Publish from + Every N days and articles are created scheduled with staggered dates on the ledger instead of the chosen status. Release due (also checked automatically at each run start) flips due posts to published - never overriding a status you changed by hand. Cron it with
php cli.php admin module=plugins/ai-assistant/generate action=releasefor hands-free drip. - Unattended jobs (cron). Three idempotent CLI actions, safe on a schedule:
…/generate action=release(publish due drip),…/generate action=refreshcron language=1 limit=5(rewrite the oldest posts),…/translate action=translatecron source=1 targets=2,3 limit=10(translate the newest missing items). CLI skips CSRF (no session exists there); every step lands in the runs ledger with acron-YYYYMMDDjob id. - Write → translate pipeline. Also translate to fans every finished article out to more languages in the same run (same provider call path as bulk translate, translations linked on the ledger so re-runs skip them).
- Grounded writing. Source URLs are fetched once per run and injected as background material (budget-capped) so articles can cite real facts instead of hallucinating them.
- Refresh stale posts. Switch to Refresh mode, Find stale posts lists candidates by age with checkboxes, Start rewrites the checked ones in place (slug untouched, honors tone + budget).
- Internal linking. After a write, Add links appends a related-reading block (keyword search over published posts, idempotent marker, skipped when nothing related).
- Social pack. Tick Social pack for an X thread, LinkedIn post and newsletter teaser per article, with copy buttons right in the result row.
- AI performance report. The writer page reports AI-written vs human posts and views over the ledger (proving whether any of this earns its keep).
- Upstream timeouts. A step failing with Upstream request failed … timed out … 0 bytes received means the provider never started answering within 180s (queue, reasoning prefill, cold model, huge prompt) - not a stalled download, so raising
max_tokenswon't help. The item fails individually and resumes on re-run; check provider status, trim grounding sources, switch reasoning off, or try a faster model before reaching for longer timeouts. - Token budgets. A per-run token cap on both bulk pages halts the stepper (stays resumable) with a notice; the server double-checks the client-reported spend.
- Writer extras. Tone presets (Settings, Name | instruction lines), per-topic model tags merged with run-wide tags, alternative headlines in each result, multi-site fan-out, CSV topic upload, multi-category, auto-categorize (the model files each article from your category list, validated server-side), QA sieve (thin/tell flags in the log), series mode (part numbering + series plan in every prompt), and a translate glossary (Original = Replacement lines in Settings) for consistent bulk translation.
- Reasoning toggles. Chain-of-thought is pure overhead on strict bulk tasks (you pay output rates for tokens nobody reads). Translate reasoning defaults off and Write reasoning stays on; either sends
reasoning_effort: none/minimalonly when the model's catalog entry declares it, only on bearer chat dialects, and silently retries without the flag if a gateway rejects it. - Product writer. The writer's Write switch covers products: descriptions with meta, disabled-by-default (publish = enable explicitly), covers, pipeline translation and social on the same runner. Scheduled drip and categories are post-only; product duplicates ledger separately.
- Refresh, SEO, links. Refresh mode rewrites stale posts in place (slug untouched); SEO backfill fills only missing excerpt/meta. Add links appends related-reading blocks found by keyword search. Every overwrite records a content revision first, so the editor's revision history is your undo.
- Review & bulk publish. Needs review lists recently written posts and products from the ledger with editor links; tick and Publish listed flips drafts/scheduled (or disabled products) live, skipping anything already published. Every bulk overwrite records a content revision first, so the editor's revision history is your undo.
- Product writer & parity. The Write switch covers products end-to-end: descriptions, disabled-by-default, covers (+gallery from extra stock hits), pipeline translation, social, refresh, SEO backfill and internal links all branch on post/product, with separate duplicate ledgers and a product performance report.
- Cost & transparency. Preview shows a dollar forecast at the active model's catalog rates; completion notices report tokens used; every bulk step lands in the Recent bulk runs ledger; an opt-in Disclose checkbox appends the configured AI note (Settings → AI disclosure text).
- Runs ledger & report. Every bulk step records calls + tokens grouped by job id (Recent bulk runs table); the performance card compares AI vs human posts/views with click-through drill-down to the actual posts; Upcoming drip lists scheduled releases (posts and products). Script tags carry mtime cache-busting so fixes actually arrive in the browser.
- Social viewer. Saved packs rehydrate per post: the result-row Social button expands the full channel texts with copy buttons, backed by the
ai_socialtable instead of the volatile run response. - Prompt playground. Settings page lab for trying system/user prompts against the active model (templates included for every job shape: chat, writer, translate, outline, SEO, social), with token usage per run. Nothing is saved.
- Product SEO preview. The stale/SEO finder and all runners honor the post/product switch, so backfill and refresh cover catalogs too. Note: scraping is against Unsplash's ToS (they prefer the API); if they block it, stock mode degrades gracefully to suggestion → generation.
- Taxonomy & menu translation. The Type dropdown on AI Translate also covers Taxonomies (category/tag names across all taxonomies) and Menus (item labels): same source/target pickers, title search, preview matrix and stepped runner, translating name (+ empty content) with glossary and JSON mode. Slugs are preserved (menus) or left for core to regenerate (categories, empty); categories carrying SEO meta in any language are skipped, because the category save rewrites every language row and cannot preserve meta.
- Translation memory. Identical source text is translated once and reused from the
ai_memorytable (keyed by normalized text + language pair + model), so repeated boilerplate, specs and headers stop costing tokens on every run. Same text on a different model still translates fresh; prune old rows manually if the table grows. - Media alt text. The writer's fourth mode describes library images into their caption field: Find images missing alt lists the first captioned-less images, Start sends each (downscaled to 1024px) to a vision model. Requires an image-capable active model (
modalitiesin the catalog) - otherwise the run errors with guidance instead of burning calls. Results log per image; failures skip individually. - Monthly token cap. Monthly token cap (Settings, 0 = off) gates every bulk action per admin and calendar month: runs fail fast with used/cap numbers once reached, and both previews show month-to-date spend next to the estimate. CLI cron counts under its own (usually zero) admin, so cap cron separately if it matters.
- Presets & webhooks. Export downloads the current filter set as JSON; Import loads it back (review, then Preview). Job webhook (Settings) POSTs
{page, processed, usage, at}fire-and-forget when a bulk job finishes - point it at your own HTTPS endpoint for n8n/Zapier handoffs. - Bulk translate (Plugins → AI Translate) translates posts and products across languages in resumable steps: filter by type, status, site, category (with children), date and title, preview the missing-translation matrix with a call estimate, then run with progress, per-item skip-and-report errors, stop/resume, and bulk JSON export. Only empty translations are touched by default. Oversized documents translate piece by piece (block-boundary chunking) instead of skipping via the old context guard. Uses the translate provider/model/key settings, so a cheap model keeps it affordable; strict tasks request JSON mode with automatic fallback.
- Sessions manager (Plugins → AI Sessions) lists every chat with title search (or Everywhere mode, which also scans message text), sorting and paging; sessions open at
#uuidlinks for sharing. Each row opens a read-only transcript each row opens a read-only transcript with the same collapsible thinking/tool cards plus a message/token/cost summary priced at the session's own model when known, and offers rename, fork (independent copy under a fresh id), pin (per-browser favorites sort first), delete (bulk included), single or bulk JSON export (with an attachments checkbox for a small text-only file) and Continue in sidebar, which hands the session to the site-editor sidebar. Sessions living only in the browser (never synced - e.g. from installs predating the server store, which self-creates on first use) are flagged with a one-click Sync to server banner instead of silently missing. - Sidebar sessions can be filtered from the dropdown itself and pinned favorites always list first.
Troubleshooting
- 401 / invalid key - re-enter the key for the active provider and save.
- 404 / no such model - the id was retired or mistyped; Refresh the list or pick from the dropdown.
- 429 / rate limit - the assistant waits and retries on its own; lower pace or raise quota for persistent limits.
- Context full - use Compact, start a new session, or lower max tokens.
- Attachment refused - the message names the supported kinds; switch to a multimodal model for images, PDFs or audio.
- Empty model list - rebuild with the Catalog button (or the CLI command above).
- 500 from the upstream gateway - almost always a dialect mismatch: a chat-shaped body (
messages/max_tokens) sent to a/responsesendpoint, or the reverse. It happens with unrouted model ids plus a stale endpoint override. The proxy now rejects these with a plain-language 400 before relaying; fix by re-selecting the model in settings (stores its route) and saving, or by pointing the override at the matching endpoint.