Skip to main content

nxt-ai-assistant

Purpose​

nxt-ai-assistant is a chat assistant for mini-grid operations. Customers and staff talk to it in Telegram (and via HTTP). It can check meters, mint or resend tokens, draft a site layout, pull KPIs, open a ticket, and run multi-step automations that operators author themselves.

The git repository and most process names still say Anansi. The admin login screen calls the product Mini-Grids Assistant.

It is a chat orchestrator with tools, not a hosted model. Gemini is the default LLM; you can point generation at OpenRouter instead. Mini-grid behavior comes from the tools, prompts, and knowledge you attach — not from a custom model.

Scope​

  • In scope:
    • Chat orchestrator (chat_orchestrator/, port 8000): conversations, Telegram webhook, scheduled jobs, skill-builder helpers. Also hosts the optional MCP gateway (mounted at /mcp-gateway) so staff can call the same tools from Claude/Codex — that gateway is not a third App Platform service.
    • Tool layer (mcp_servers/): MCP servers for meters, grid design, Jira, Grafana, payments, solar, knowledge, and more. Most servers are off until you set {SERVER}_ENABLED.
    • Prompt library (shared/prompts/): bundled defaults, in-app edits, optional Google Doc per prompt, plus context modules (built-in, curated, or an attached Drive file) that can pin to a prompt or a skill.
    • RAG ingestion (rag_pipeline/): GitHub, Google Drive, Telegram, Grafana — usually as a scheduled job inside the orchestrator, not a third process.
    • Admin UI (anansi_app/, NiceGUI, port 8501): history, prompts, settings, Skills, broadcasts, grid design, and an in-app chat widget.
    • Customer chat widget (mini_app/).
  • Out of scope:
    • Hosting the LLM yourself (calls Gemini or OpenRouter).
    • WhatsApp (not supported yet; Telegram is the messenger surface).
    • Being your metering or billing system — those stay in nxt-backend / nxt-device-messaging (or whatever API you wire through MCP).
    • Writing to the auth database. This app reads operational records there; its own tables and migrations live in the chat database.

What this system does in production​

  • Accepts a message on POST /chat (or POST /) from Telegram or an API caller, loads history, resolves prompts, calls the LLM, runs tools if asked, and replies.
  • Treats staff vs customer by organization_id vs STAFF_ORG_ID (default 2): staff get the staff prompt and full tools; customers get a narrower support prompt.
  • Lets operators edit prompts in the admin app without a redeploy (about a minute to take effect). Bundled files always remain the fallback. A Google Doc is optional per prompt.
  • Escalates conversations that need a human: Jira when it is configured and healthy, otherwise an internal ticket ledger so routing still works.
  • Runs operator-built Skills (chat-authored steps, schedulable) and hardcoded expert flows such as /lpp (site package), /analyze / /kpi / /report.
  • Optionally exposes a per-user MCP endpoint at /mcp-gateway/mcp (Google sign-in, no shared API key). Equipment control, payments, and messaging tools stay off that endpoint even when they are on for the bot.

Production App Platform is two services (orchestrator + admin). Local Compose still runs orchestrator and a separate tools HTTP bridge.

Keep one replica of the orchestrator: in-flight Telegram workflows, alert correlation locks, and similar state live in that process.

Primary workflows​

  • Customer or staff chat: Telegram (secret token) or X-Api-Key → orchestrator → prompt resolve (DB override → optional Doc → bundled file) plus any attached context modules → LLM → optional tool calls → persist in the chat database → reply (HTTP body for API; Telegram Bot API for the bot).
  • Admin chat widget: signed-in operator in anansi-app talks over HTTP (source=web) with page context (grid design, tickets). Same orchestrator as Telegram.
  • Site package (/lpp): GPS + site name → GRID3 boundary (if you have the GeoPackages) → layout engine → Google Doc. Needs Google credentials; GRID3 files are optional.
  • Skill: operator authors steps in the admin builder → validate/summarize via orchestrator → save → run now or on a schedule with nobody in the loop. Skills can pin the same context modules as prompts.
  • External MCP client: claude mcp add --transport http … /mcp-gateway/mcp → Google OAuth in the browser → tools filtered to that user’s org. Unset MCP_GATEWAY_BASE_URL or MCP_GATEWAY_TOKEN_SECRET and the gateway is never mounted; chat is unchanged.
  • Knowledge refresh: APScheduler in the orchestrator runs RAG ingest (GitHub / Drive / Grafana) when those flags are on. Manual: python rag_pipeline/ingestion/batch_ingest.py inside that service.

Setup and run​

  • Repository: github.com/nxtgrid/nxt-ai-assistant

  • Python 3.11+. Gemini API key for the default provider. Chat database (Supabase) plus a separate read-only auth/ops database. Google service account if you use Docs/Drive.

  • Shared package: ./setup_shared.sh and PYTHONPATH (see README). Per-service venvs are documented there too.

  • Local Compose (orchestrator 8000, tools 8080):

    docker compose up

    TOOLS_SERVICE_URL=http://tools-service:8080 is set for the orchestrator container.

  • Health: GET http://localhost:8000/health → status: healthy. Admin: GET http://localhost:8501/healthz. Gateway (if mounted): GET /mcp-gateway/healthz.

Full env catalog: chat_orchestrator/.env.example, mcp_servers/.env.example, shared/config/flag_registry.py. Settings in the admin app are the usual way to change operator flags; credentials stay in the platform env and are not written back.

Chat-database migrations live in db/migrations/ and are not applied by deploy. An unapplied migration typically fails at runtime (UndefinedTableError), not at build time.

Deployment​

  • DigitalOcean App Platform — two components, one replica each; example spec .do/app.example.yaml. Ingress path prefixes matter: a new public path that is not listed falls through to the admin login page.

APIs and interfaces​

Orchestrator (chat_orchestrator/orchestrator/api/app.py):

MethodPathAuthRole
GET/healthnoneLiveness (status: healthy) plus a ticket-correlation failure count.
POST/ or /chatX-Api-Key or Telegram secretMain chat. API key: JSON body response. Telegram: reply via Bot API.
POST/skills/validate, /skills/summarize, /skills/dispatch-scheduleAPI keySkill builder (admin).
POST/chat/notifyX-Notify-SecretExternal alerts (n8n / VRM / Grafana). Fail-closed if the secret is unset.
POST/webhook/jira(Jira webhook)Ticket events. Must reach this service (/webhook prefix), not the admin UI.
GET/api/v1/jobsX-Api-KeyAPScheduler job list.

MCP gateway (mounted when both gateway env vars are set). Paths below are after the /mcp-gateway prefix:

PathAuthRole
/mcpBearer (OAuth)Streamable HTTP MCP. No token → 401 + WWW-Authenticate.
/healthznoneConfirms the mount succeeded.
/oauth/register, /oauth/authorize, /oauth/google-callback, /oauth/token(OAuth)Dynamic client registration, Google sign-in, PKCE token exchange.

Discovery documents for that issuer also live at /.well-known/oauth-authorization-server/mcp-gateway and /.well-known/oauth-protected-resource/mcp-gateway on the orchestrator root (they are not under /mcp-gateway).

IDENTITY_ASSERTION_KEY is not API_KEY. Only callers that hold it may assert user_email when the auth DB lookup misses (skill builder). Unset = nobody gets that fallback.

Local tools bridge (mcp_servers/bridge.py, port 8080): X-API-Key on /servers, tool list, and POST /servers/{name}/tools/{tool}. Errors are sanitized before they go back to the model. In production on App Platform, tools are imported in-process by the orchestrator, not this extra HTTP service.

Integrations and dependencies​

  • LLM: Gemini by default; LLM_PROVIDER=openrouter for OpenAI-style completions (shared/llm/).
  • Data: Chat database (writeable; conversations, prompts, skills, RAG, OAuth code single-use). Auth/ops database (read-only; accounts, grids, organizations). Optional Timescale / Victron VRM when analytics tools are enabled.
  • Google: Docs + Drive for optional prompt docs, context modules, and /lpp output. The same Web OAuth client is reused by the admin login and the MCP gateway (add a second authorized redirect URI for the gateway callback).
  • Messenger: Telegram bot + webhook. Mini App for operator-portal chat.
  • Optional tools: Metering API, Jira, Grafana, GRID3 GeoPackages, Langfuse traces (LANGFUSE_ENABLED).
  • License: MPL-2.0.

Operations notes​

  • Two kinds of config: platform secrets (never edited from the Settings UI into the bot process) vs flags in flag_registry.py (Settings page; on DigitalOcean, saving can patch the live app spec and redeploy).
  • Prompt resolve order: DB override → attached Doc → bundled shared/prompts/library/*.prompt. Cache: ~1 minute for DB, up to an hour for Docs; “Reload cache” on the prompt detail.
  • Ingress is last-match / catch-all: /chat, /mini-app, /api/mini-app, /webhook, /mcp-gateway, and the two /.well-known/oauth-*/mcp-gateway paths must go to the orchestrator with the path prefix kept. Everything else is the admin UI.
  • Change impact:
    • if chat/tool HTTP shapes change, Telegram, n8n, the admin skill builder, and MCP clients all need a look;
    • if STAFF_ORG_ID is wrong, customers get staff tools (or staff get the customer prompt);
    • if /webhook is routed to anansi-app instead of the orchestrator, Jira callbacks look successful and nothing happens in chat;
    • if MCP_GATEWAY_BASE_URL does not end in /mcp-gateway, discovery and the Google redirect disagree with the mount path.
  • Failure modes:
    • missing Gemini/OpenRouter key → no generations;
    • tools service down (local Compose) or {SERVER}_ENABLED unset → tool calls fail, chat may still answer;
    • Google Doc unreachable → bundled/DB prompt still works;
    • API_KEY unset on the tools bridge → 500, not 401;
    • second orchestrator replica → two processes can resume the same in-flight Telegram workflow, or file duplicate tickets for the same grid alert;
    • Chat DB migration 0032_oauth_code_single_use.sql not applied → MCP token exchange fails at the last step;
    • Auth DB treated as writable / chat migrations pointed at it → those writes cannot work.

Source of truth​

  • Repository: github.com/nxtgrid/nxt-ai-assistant
  • Product and setup: README.md, llms.txt
  • Orchestrator: chat_orchestrator/orchestrator/api/app.py, chat_orchestrator/.env.example
  • Tools: mcp_servers/bridge.py, mcp_servers/servers/
  • MCP gateway: mcp_servers/gateway/ (tiers.py is the allowlist)
  • Prompts: shared/prompts/, shared/config/flag_registry.py
  • Admin: anansi_app/
  • RAG: rag_pipeline/
  • Chat DB migrations: db/migrations/
  • App spec: .do/app.example.yaml