Files
llm-in-text/CLAUDE.md
T

6.1 KiB

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with this repository.

Project version: v0.2.0 (full-stack architecture upgrade since b82c6d3)

Project Overview

LLM in Text is an AI-powered Markdown editor built with Vue 3 + Vite (frontend) and FastAPI + Python + Ollama + Redis Streams (backend). It provides real-time AI completion suggestions, OCR image recognition, document conversion (PDF/DOCX/PPTX to Markdown), TTS text-to-speech, web search integration (via SearXNG + Firecrawl), CAPTCHA verification, and async job queue processing.

  • Completion interface uses plain POST/JSON (not SSE). Frontend sends X-Request-Id and calls /v1/completions/cancel on abort.
  • AI completion is disabled when the document exceeds 32 KB (enforced both in UI and plugin layer).
  • OCR text and doc-block content are injected into completion context as hidden context — they should NOT be rendered as visible user text.
  • /v1/convert supports txt, docx, pptx, pdf via MarkItDown. Non-txt files go through Markdown sanitization (image removal, newline compression).
  • /v1/export/pdf is called from the frontend but may not be implemented on the backend — verify before debugging PDF export.
  • Task queue architecture (v0.2.0 new): Backend shifted from synchronous endpoints to Redis Streams async queue. job_system.py defines JOB_TYPES (completion/pro_completion/web_search/compress/ocr/convert/tts/asr), worker.py consumes the queue, job_handlers.py registers handlers per type. All tasks support concurrency control, rate limiting, and circuit breakers.
  • Session tracking: session_store.py provides InMemorySessionStore (dev) and PostgresSessionStore (prod), tracking request identity via session_hash + ip_hash.
  • Risk control system: risk_config.py + risk_control.py implement rate limiting (sliding window), concurrency limits, circuit breaker pattern, and budget tracking — all thresholds via environment variables.

Quick Start

# Frontend
npm install
npm run dev        # Vite dev server on port 5173, proxies /v1 to backend

# Backend
pip install -r backend/requirements.txt
python backend/main.py                    # port 8001

# Tests (90% coverage gate on backend modules)
pytest                                    # full suite with coverage

# Single test file (faster, no coverage overhead)
pytest backend/tests/test_prompt.py -v --no-cov

# Build for production
npm run build

Architecture

Frontend (src/)

Layer Key Files Responsibility
Entry main.js, App.vue Vue app bootstrap, Pinia + Router mount
Routing router/index.js / → EditorView, /docs → DocsView
Editor components/MilkdownEditor.vue Central control: Crepe editor, plugin registration, upload/export/OCR/TTS/AI toggle, 32 KB limit
Plugins (TypeScript) plugins/copilotPlugin.ts — ghost text, request scheduling, cancel, language detection, hidden context injection
Plugins (TypeScript) plugins/docBlockPlugin.ts — doc-block nodes and rendering
Plugins (TypeScript) plugins/mermaidPlugin.ts — Mermaid diagram preview
Store stores/settings.js localStorage-persisted settings (theme, modelThinking, debounceMs, privacyMode, language, background*, ttsInstruct)
API utils/api.js — fetchSuggestion, cancel completion, TTS requests; config.js — VITE_* env-based URL config
Utilities utils/convert.js, ocrCache.js, docBlock.js, i18n.js

Backend (backend/)

File Responsibility
main.py FastAPI app, CORS, API key auth, routes: /v1/completions, /v1/ocr, /v1/convert, /v1/completions/cancel. TTS routes lazily registered from tts_asr.py.
llm.py Async Ollama calls (call_ollama, stream_ollama) and VLM OCR (call_vlm_ocr). Timeout control.
prompt.py Prompt assembly: build_completion_prompts, prepare_prompt_context. Templates from prompts/ directory.
pro_completions.py Pro-tier completion endpoint (newer addition).
tts_asr.py TTS text-to-speech. Late-registered routes via _register_tts_asr_routes.
geoip.py Client IP location lookup for non-privacy-mode requests.

Request Flow: Completion

MilkdownEditor.vue → copilotPlugin.ts (debounce, abort, language detection)
  → utils/api.js (fetchSuggestion: generates request_id, AbortSignal, reads settings)
    → backend/main.py (/v1/completions: auth, prompt context, call_ollama via asyncio.Task)
      → backend/prompt.py (system + user prompt from prefix/suffix/context)
        → backend/llm.py (call_ollama to Ollama)
    ← JSON { content, request_id }
  → copilotPlugin.ts (insertGhostText into editor)

Debugging Paths

Issue Trace Order
Completion not firing MilkdownEditor.vuecopilotPlugin.ts (check enabled, size limit, debounce)
Wrong completion result prompt.pyllm.py. Check prompt context and language detection.
Cancel not working main.py request_id lifecycle ↔ frontend X-Request-Id + cancel call
OCR empty result main.py base64 decode → llm.py call_vlm_ocr
Document conversion dirty _sanitize_converted_markdown in main.py

Naming Conventions (Mixed)

  • Vue components/views: PascalCase (MilkdownEditor.vue)
  • Frontend utils/config: lowercase .js (api.js, config.js)
  • Plugin layer: TypeScript (.ts)
  • Python backend: snake_case

Follow the style of each file. Do not reformat across directories for consistency. UI copy defaults to Chinese.

Important Rules

  • Do not modify milkdown-docs/ (read-only reference).
  • Code and tests override README.md when they conflict — the README is partially outdated.
  • Plugin code (copilotPlugin.ts) is state-machine-style: small changes can break subtle interactions. Change one thing at a time and verify in-browser.
  • No hardcoded secrets, empty catch/except blocks, as any, or @ts-ignore in new code.
  • Subdirectory AGENTS.md files contain more detailed guidance: ./AGENTS.md (root), backend/AGENTS.md, src/AGENTS.md, src/plugins/AGENTS.md. Read them when working in those areas.