6.1 KiB
6.1 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with this repository.
Project version: v0.2.0 (full-stack architecture upgrade since b82c6d3)
Project Overview
LLM in Text is an AI-powered Markdown editor built with Vue 3 + Vite (frontend) and FastAPI + Python + Ollama + Redis Streams (backend). It provides real-time AI completion suggestions, OCR image recognition, document conversion (PDF/DOCX/PPTX to Markdown), TTS text-to-speech, web search integration (via SearXNG + Firecrawl), CAPTCHA verification, and async job queue processing.
- Completion interface uses plain POST/JSON (not SSE). Frontend sends
X-Request-Idand calls/v1/completions/cancelon abort. - AI completion is disabled when the document exceeds 32 KB (enforced both in UI and plugin layer).
- OCR text and doc-block content are injected into completion context as hidden context — they should NOT be rendered as visible user text.
/v1/convertsupports txt, docx, pptx, pdf via MarkItDown. Non-txt files go through Markdown sanitization (image removal, newline compression)./v1/export/pdfis called from the frontend but may not be implemented on the backend — verify before debugging PDF export.- Task queue architecture (v0.2.0 new): Backend shifted from synchronous endpoints to Redis Streams async queue.
job_system.pydefines JOB_TYPES (completion/pro_completion/web_search/compress/ocr/convert/tts/asr),worker.pyconsumes the queue,job_handlers.pyregisters handlers per type. All tasks support concurrency control, rate limiting, and circuit breakers. - Session tracking:
session_store.pyprovides InMemorySessionStore (dev) and PostgresSessionStore (prod), tracking request identity via session_hash + ip_hash. - Risk control system:
risk_config.py+risk_control.pyimplement rate limiting (sliding window), concurrency limits, circuit breaker pattern, and budget tracking — all thresholds via environment variables.
Quick Start
# Frontend
npm install
npm run dev # Vite dev server on port 5173, proxies /v1 to backend
# Backend
pip install -r backend/requirements.txt
python backend/main.py # port 8001
# Tests (90% coverage gate on backend modules)
pytest # full suite with coverage
# Single test file (faster, no coverage overhead)
pytest backend/tests/test_prompt.py -v --no-cov
# Build for production
npm run build
Architecture
Frontend (src/)
| Layer | Key Files | Responsibility |
|---|---|---|
| Entry | main.js, App.vue |
Vue app bootstrap, Pinia + Router mount |
| Routing | router/index.js |
/ → EditorView, /docs → DocsView |
| Editor | components/MilkdownEditor.vue |
Central control: Crepe editor, plugin registration, upload/export/OCR/TTS/AI toggle, 32 KB limit |
| Plugins (TypeScript) | plugins/copilotPlugin.ts — ghost text, request scheduling, cancel, language detection, hidden context injection |
|
| Plugins (TypeScript) | plugins/docBlockPlugin.ts — doc-block nodes and rendering |
|
| Plugins (TypeScript) | plugins/mermaidPlugin.ts — Mermaid diagram preview |
|
| Store | stores/settings.js |
localStorage-persisted settings (theme, modelThinking, debounceMs, privacyMode, language, background*, ttsInstruct) |
| API | utils/api.js — fetchSuggestion, cancel completion, TTS requests; config.js — VITE_* env-based URL config |
|
| Utilities | utils/convert.js, ocrCache.js, docBlock.js, i18n.js |
Backend (backend/)
| File | Responsibility |
|---|---|
main.py |
FastAPI app, CORS, API key auth, routes: /v1/completions, /v1/ocr, /v1/convert, /v1/completions/cancel. TTS routes lazily registered from tts_asr.py. |
llm.py |
Async Ollama calls (call_ollama, stream_ollama) and VLM OCR (call_vlm_ocr). Timeout control. |
prompt.py |
Prompt assembly: build_completion_prompts, prepare_prompt_context. Templates from prompts/ directory. |
pro_completions.py |
Pro-tier completion endpoint (newer addition). |
tts_asr.py |
TTS text-to-speech. Late-registered routes via _register_tts_asr_routes. |
geoip.py |
Client IP location lookup for non-privacy-mode requests. |
Request Flow: Completion
MilkdownEditor.vue → copilotPlugin.ts (debounce, abort, language detection)
→ utils/api.js (fetchSuggestion: generates request_id, AbortSignal, reads settings)
→ backend/main.py (/v1/completions: auth, prompt context, call_ollama via asyncio.Task)
→ backend/prompt.py (system + user prompt from prefix/suffix/context)
→ backend/llm.py (call_ollama to Ollama)
← JSON { content, request_id }
→ copilotPlugin.ts (insertGhostText into editor)
Debugging Paths
| Issue | Trace Order |
|---|---|
| Completion not firing | MilkdownEditor.vue → copilotPlugin.ts (check enabled, size limit, debounce) |
| Wrong completion result | prompt.py → llm.py. Check prompt context and language detection. |
| Cancel not working | main.py request_id lifecycle ↔ frontend X-Request-Id + cancel call |
| OCR empty result | main.py base64 decode → llm.py call_vlm_ocr |
| Document conversion dirty | _sanitize_converted_markdown in main.py |
Naming Conventions (Mixed)
- Vue components/views: PascalCase (
MilkdownEditor.vue) - Frontend utils/config: lowercase
.js(api.js,config.js) - Plugin layer: TypeScript (
.ts) - Python backend: snake_case
Follow the style of each file. Do not reformat across directories for consistency. UI copy defaults to Chinese.
Important Rules
- Do not modify
milkdown-docs/(read-only reference). - Code and tests override README.md when they conflict — the README is partially outdated.
- Plugin code (
copilotPlugin.ts) is state-machine-style: small changes can break subtle interactions. Change one thing at a time and verify in-browser. - No hardcoded secrets, empty catch/except blocks,
as any, or@ts-ignorein new code. - Subdirectory AGENTS.md files contain more detailed guidance:
./AGENTS.md(root),backend/AGENTS.md,src/AGENTS.md,src/plugins/AGENTS.md. Read them when working in those areas.