Compare commits
4 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 9904b9bd78 | |||
| 7ed199aaf1 | |||
| 9ff51ac2f3 | |||
| be4000b774 |
+7
-1
@@ -1,4 +1,4 @@
|
|||||||
# Logs
|
# Logs
|
||||||
logs
|
logs
|
||||||
*.log
|
*.log
|
||||||
npm-debug.log*
|
npm-debug.log*
|
||||||
@@ -39,3 +39,9 @@ env/
|
|||||||
*.njsproj
|
*.njsproj
|
||||||
*.sln
|
*.sln
|
||||||
*.sw?
|
*.sw?
|
||||||
|
|
||||||
|
|
||||||
|
# IDE directories
|
||||||
|
.kilocode/
|
||||||
|
.kilo/
|
||||||
|
.codex/
|
||||||
@@ -0,0 +1,52 @@
|
|||||||
|
# 导出按钮缺失修复计划
|
||||||
|
|
||||||
|
## 问题分析
|
||||||
|
当前 `action-buttons` 区域只有以下按钮可见:
|
||||||
|
- 上传文件
|
||||||
|
- 导入 Markdown
|
||||||
|
- 导出 Markdown
|
||||||
|
- 上传图片
|
||||||
|
- AI 切换按钮
|
||||||
|
|
||||||
|
**缺失功能**:DOCX 和 PDF 导出按钮
|
||||||
|
|
||||||
|
## 调查结果
|
||||||
|
1. ✅ 翻译文件中已存在 `exportDocx` 和 `exportPdf` 键名(src/utils/i18n.js)
|
||||||
|
2. ❌ 模板中**完全缺失**这两个按钮的 HTML 代码
|
||||||
|
3. ❓ 导出功能后端已实现,前端只需要添加调用接口的按钮
|
||||||
|
4. ✅ 相关 CSS 样式已存在,按钮外观无需额外调整
|
||||||
|
|
||||||
|
## 实施计划
|
||||||
|
|
||||||
|
### 1. 添加 UI 按钮
|
||||||
|
在 `src/components/MilkdownEditor.vue:79` 之后添加两个新按钮:
|
||||||
|
- DOCX 导出按钮
|
||||||
|
- PDF 导出按钮
|
||||||
|
|
||||||
|
按钮位置:
|
||||||
|
```
|
||||||
|
导出 Markdown → 导出 DOCX → 导出 PDF → 上传图片
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2. 实现前端导出功能
|
||||||
|
使用已安装的依赖库:
|
||||||
|
- `docx` 库:用于 DOCX 导出
|
||||||
|
- `html2pdf.js` 库:用于 PDF 导出
|
||||||
|
|
||||||
|
需要添加的函数:
|
||||||
|
```javascript
|
||||||
|
const exportDocx = async () => {
|
||||||
|
// 使用 docx 库实现导出
|
||||||
|
}
|
||||||
|
|
||||||
|
const exportPdf = async () => {
|
||||||
|
// 使用 html2pdf.js 实现导出
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3. 按钮图标
|
||||||
|
- DOCX:使用文档图标
|
||||||
|
- PDF:使用 PDF 专用图标
|
||||||
|
|
||||||
|
### 4. 状态管理
|
||||||
|
添加加载状态和错误处理,与现有按钮保持一致风格
|
||||||
@@ -1,4 +1,19 @@
|
|||||||
# 仓库指南
|
# rules.md
|
||||||
|
|
||||||
|
在构建这个LLM应用网页时,你需要基于VUE3开发。我需要前端只运行渲染和数据回传,后端负责llm api调用,类似copilet的auto inline suggustions实现和数据解析。
|
||||||
|
|
||||||
|
# **重要** : 在回复用户消息时,一定要使用中文
|
||||||
|
|
||||||
|
## 指导原则
|
||||||
|
|
||||||
|
- 不要擅自用npm或者yarn运行网页,你既看不到网页的内容,也无法阻止命令暂停。但是,你可以用npm run build检查代码。
|
||||||
|
- 应该保证代码效率,不多定义变量,不写冗余注释,把降低延迟放在第一位。
|
||||||
|
- 每次完成任务前都要反复阅读检查代码,确保代码准确无误。
|
||||||
|
- 尽量不要搜索关键字,而是了解代码结构后查询整个问题代码明确问题所在。
|
||||||
|
- @/milkdown-docs/ 代表milkdown的最新官方文档,不要修改,涉及到前端编辑器的指令时要核对官方文档。
|
||||||
|
|
||||||
|
|
||||||
|
# 仓库指南
|
||||||
|
|
||||||
## 语言约定
|
## 语言约定
|
||||||
项目文档、日志、错误提示以及对外返回的文字信息统一使用 **中文**。前端 UI 默认展示中文,若需多语言支持请在相应模块实现。
|
项目文档、日志、错误提示以及对外返回的文字信息统一使用 **中文**。前端 UI 默认展示中文,若需多语言支持请在相应模块实现。
|
||||||
@@ -48,5 +63,3 @@ dist/ # 构建产出(生成文件)
|
|||||||
- 按照 `backend/main.py` 中的实现,对上传文件的大小和类型进行校验,防止滥用。
|
- 按照 `backend/main.py` 中的实现,对上传文件的大小和类型进行校验,防止滥用。
|
||||||
- 定期审计依赖安全(`npm audit`、`pip-audit`)。
|
- 定期审计依赖安全(`npm audit`、`pip-audit`)。
|
||||||
|
|
||||||
---
|
|
||||||
以上指南旨在保持贡献一致性并维护代码库健康,欢迎通过 Pull Request 提出改进。
|
|
||||||
|
|||||||
@@ -1,814 +0,0 @@
|
|||||||
# llm-in-text 修复清单(匿名可用版)
|
|
||||||
|
|
||||||
## 说明
|
|
||||||
|
|
||||||
这不是审计报告。
|
|
||||||
|
|
||||||
这份文档只回答三件事:
|
|
||||||
|
|
||||||
1. 现在具体哪里有问题
|
|
||||||
2. 问题为什么会发生
|
|
||||||
3. 应该怎么改
|
|
||||||
|
|
||||||
前提按你的要求处理:
|
|
||||||
|
|
||||||
- 网站是匿名可用的
|
|
||||||
- 不做用户登录
|
|
||||||
- 不做用户身份体系
|
|
||||||
- 但仍然要防止接口被滥用、站点被刷爆、服务被恶意调用
|
|
||||||
|
|
||||||
匿名可用不等于完全不做保护。
|
|
||||||
|
|
||||||
对于这种网站,正确做法通常是:
|
|
||||||
|
|
||||||
- 不做用户登录
|
|
||||||
- 不在前端放任何真正的服务端秘密
|
|
||||||
- 用服务端限流、来源限制、请求大小限制、网关策略保护接口
|
|
||||||
- 必要时用站点级防刷手段,而不是用户级登录
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 1. 前端硬编码了服务端 API Key
|
|
||||||
|
|
||||||
### 具体问题
|
|
||||||
|
|
||||||
- [src/utils/api.js:4](/C:/Users/ydy/Desktop/llm-in-text/src/utils/api.js#L4)
|
|
||||||
- [src/utils/convert.js:3](/C:/Users/ydy/Desktop/llm-in-text/src/utils/convert.js#L3)
|
|
||||||
|
|
||||||
代码里把:
|
|
||||||
|
|
||||||
```js
|
|
||||||
const API_KEY = 'your-secret-key-here'
|
|
||||||
```
|
|
||||||
|
|
||||||
直接写进了前端源码。
|
|
||||||
|
|
||||||
### 错误原因
|
|
||||||
|
|
||||||
前端代码最终会发到浏览器里。
|
|
||||||
|
|
||||||
只要用户能打开网站,就一定能在浏览器开发者工具、打包产物、网络请求里看到这个 key。
|
|
||||||
所以前端里的“密钥”根本不是密钥,只是公开字符串。
|
|
||||||
|
|
||||||
### 会导致什么
|
|
||||||
|
|
||||||
- 任何人都可以绕过你的网站,直接写脚本刷你的后端
|
|
||||||
- 这个 key 一旦被复制,就等于后端公开可调用
|
|
||||||
|
|
||||||
### 整改方式
|
|
||||||
|
|
||||||
你的场景不做登录,所以最简单、正确的做法是:
|
|
||||||
|
|
||||||
1. 删除前端里的 `API_KEY`
|
|
||||||
2. 后端不要再要求前端传固定共享 key
|
|
||||||
3. 改成下面这套匿名保护方案:
|
|
||||||
- 只允许来自你站点域名的浏览器请求
|
|
||||||
- 网关层限流
|
|
||||||
- 接口级限流
|
|
||||||
- 请求体大小限制
|
|
||||||
- 必要时加站点级验证码或 challenge,而不是登录
|
|
||||||
|
|
||||||
### 你应该改成什么
|
|
||||||
|
|
||||||
- `src/utils/api.js` 不再发 `X-API-Key`
|
|
||||||
- `src/utils/convert.js` 不再发 `X-API-Key`
|
|
||||||
- `backend/main.py` 删除固定 `API_KEY` 和对应校验逻辑
|
|
||||||
|
|
||||||
### 验收标准
|
|
||||||
|
|
||||||
- 全仓库搜不到 `your-secret-key-here`
|
|
||||||
- 前端请求头中不再包含固定共享 key
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 2. 后端 CORS 过宽
|
|
||||||
|
|
||||||
### 具体问题
|
|
||||||
|
|
||||||
- [backend/main.py:34](/C:/Users/ydy/Desktop/llm-in-text/backend/main.py#L34)
|
|
||||||
- [backend/main.py:35](/C:/Users/ydy/Desktop/llm-in-text/backend/main.py#L35)
|
|
||||||
- [backend/main.py:36](/C:/Users/ydy/Desktop/llm-in-text/backend/main.py#L36)
|
|
||||||
- [backend/main.py:37](/C:/Users/ydy/Desktop/llm-in-text/backend/main.py#L37)
|
|
||||||
|
|
||||||
现在配置是:
|
|
||||||
|
|
||||||
- `allow_origins=["*"]`
|
|
||||||
- `allow_credentials=True`
|
|
||||||
- `allow_methods=["*"]`
|
|
||||||
- `allow_headers=["*"...]`
|
|
||||||
|
|
||||||
### 错误原因
|
|
||||||
|
|
||||||
这是开发期常见的“先全开让它跑起来”的写法。
|
|
||||||
但生产里这样做会让任何站点都能更容易发起跨域调用。
|
|
||||||
|
|
||||||
### 会导致什么
|
|
||||||
|
|
||||||
- 其他网站更容易借你的浏览器接口能力
|
|
||||||
- 以后一旦加 cookie、session、任何凭据,会立刻放大风险
|
|
||||||
|
|
||||||
### 整改方式
|
|
||||||
|
|
||||||
既然你是匿名站点,不做登录,那就更应该把跨域收紧:
|
|
||||||
|
|
||||||
1. 只允许你的正式域名和本地开发域名
|
|
||||||
2. 不要开 `allow_credentials=True`,匿名站一般不需要
|
|
||||||
3. 只开放需要的方法和头
|
|
||||||
|
|
||||||
### 建议改法
|
|
||||||
|
|
||||||
把:
|
|
||||||
|
|
||||||
```python
|
|
||||||
allow_origins=["*"]
|
|
||||||
allow_credentials=True
|
|
||||||
allow_methods=["*"]
|
|
||||||
allow_headers=["*", "X-API-Key", "X-Client-IP", "X-Request-Id"]
|
|
||||||
```
|
|
||||||
|
|
||||||
改成类似:
|
|
||||||
|
|
||||||
```python
|
|
||||||
allow_origins=[
|
|
||||||
"https://your-domain.com",
|
|
||||||
"https://www.your-domain.com",
|
|
||||||
"http://localhost:5173",
|
|
||||||
]
|
|
||||||
allow_credentials=False
|
|
||||||
allow_methods=["POST", "OPTIONS"]
|
|
||||||
allow_headers=["Content-Type", "X-Request-Id"]
|
|
||||||
```
|
|
||||||
|
|
||||||
### 验收标准
|
|
||||||
|
|
||||||
- 非你自己域名的网页无法直接跨域调用你的接口
|
|
||||||
- 不再开放无用头和无用方法
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 3. 默认会去拿用户公网 IP,并发送给后端
|
|
||||||
|
|
||||||
### 具体问题
|
|
||||||
|
|
||||||
- [src/stores/settings.js:16](/C:/Users/ydy/Desktop/llm-in-text/src/stores/settings.js#L16)
|
|
||||||
- [src/utils/api.js:54](/C:/Users/ydy/Desktop/llm-in-text/src/utils/api.js#L54)
|
|
||||||
- [src/utils/api.js:100](/C:/Users/ydy/Desktop/llm-in-text/src/utils/api.js#L100)
|
|
||||||
- [backend/main.py:110](/C:/Users/ydy/Desktop/llm-in-text/backend/main.py#L110)
|
|
||||||
|
|
||||||
流程是:
|
|
||||||
|
|
||||||
1. 前端默认 `privacyMode = false`
|
|
||||||
2. 前端请求 `https://api.ipify.org?format=json`
|
|
||||||
3. 获取公网 IP
|
|
||||||
4. 放进 `X-Client-IP`
|
|
||||||
5. 后端再做地理位置推断
|
|
||||||
|
|
||||||
### 错误原因
|
|
||||||
|
|
||||||
这是把“个性化上下文”做成了默认行为。
|
|
||||||
但对匿名站点来说,这不是必要信息。
|
|
||||||
|
|
||||||
### 会导致什么
|
|
||||||
|
|
||||||
- 页面会额外访问第三方服务
|
|
||||||
- 用户 IP 会进入你的请求链路
|
|
||||||
- 模型上下文中会混入地理位置信息
|
|
||||||
|
|
||||||
### 整改方式
|
|
||||||
|
|
||||||
如果你不需要真正的按地理位置个性化,就最简单:
|
|
||||||
|
|
||||||
1. 删除 `getClientIP()`
|
|
||||||
2. 删除调用 `api.ipify.org`
|
|
||||||
3. 删除 `X-Client-IP`
|
|
||||||
4. 后端删除 GeoIP 逻辑
|
|
||||||
5. `privacyMode` 可以保留,但默认应是更安全的行为
|
|
||||||
|
|
||||||
### 你应该删什么
|
|
||||||
|
|
||||||
- `src/utils/api.js` 中的 `getClientIP`
|
|
||||||
- `headers['X-Client-IP'] = clientIP`
|
|
||||||
- `backend/main.py` 中 `get_client_ip`
|
|
||||||
- `location = get_ip_location_text(client_ip)`
|
|
||||||
- `geoip.py` 如果以后不用可以移除
|
|
||||||
|
|
||||||
### 验收标准
|
|
||||||
|
|
||||||
- 前端网络面板中不再出现 `api.ipify.org`
|
|
||||||
- 后端不再接收 `X-Client-IP`
|
|
||||||
- prompt 不再包含用户位置
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 4. 后端把内部异常原样返回给前端
|
|
||||||
|
|
||||||
### 具体问题
|
|
||||||
|
|
||||||
- [backend/main.py:184](/C:/Users/ydy/Desktop/llm-in-text/backend/main.py#L184)
|
|
||||||
- [backend/main.py:251](/C:/Users/ydy/Desktop/llm-in-text/backend/main.py#L251)
|
|
||||||
- [backend/main.py:303](/C:/Users/ydy/Desktop/llm-in-text/backend/main.py#L303)
|
|
||||||
|
|
||||||
现在写法是:
|
|
||||||
|
|
||||||
```python
|
|
||||||
return JSONResponse(content={"error": str(e)}, status_code=500)
|
|
||||||
```
|
|
||||||
|
|
||||||
### 错误原因
|
|
||||||
|
|
||||||
这是开发期为了调试方便常见的写法。
|
|
||||||
但线上不应该把真实异常直接发给浏览器。
|
|
||||||
|
|
||||||
### 会导致什么
|
|
||||||
|
|
||||||
- 暴露内部实现细节
|
|
||||||
- 暴露依赖报错、路径、上游信息
|
|
||||||
|
|
||||||
### 整改方式
|
|
||||||
|
|
||||||
统一改成:
|
|
||||||
|
|
||||||
1. 前端只收到固定错误码和通用提示
|
|
||||||
2. 后端日志里保留详细异常
|
|
||||||
3. 返回 request id 方便排查
|
|
||||||
|
|
||||||
### 建议响应格式
|
|
||||||
|
|
||||||
```json
|
|
||||||
{
|
|
||||||
"error": {
|
|
||||||
"code": "UPSTREAM_TIMEOUT",
|
|
||||||
"message": "Service temporarily unavailable",
|
|
||||||
"request_id": "xxxx"
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### 验收标准
|
|
||||||
|
|
||||||
- 前端不再收到 Python 原始报错
|
|
||||||
- 日志可通过 request id 查到真实错误
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 5. `/v1/convert` 和 `/v1/ocr` 没有文件安全边界
|
|
||||||
|
|
||||||
### 具体问题
|
|
||||||
|
|
||||||
- [backend/main.py:239](/C:/Users/ydy/Desktop/llm-in-text/backend/main.py#L239)
|
|
||||||
- [backend/main.py:268](/C:/Users/ydy/Desktop/llm-in-text/backend/main.py#L268)
|
|
||||||
- [backend/main.py:275](/C:/Users/ydy/Desktop/llm-in-text/backend/main.py#L275)
|
|
||||||
- [backend/main.py:281](/C:/Users/ydy/Desktop/llm-in-text/backend/main.py#L281)
|
|
||||||
- [src/utils/convert.js:22](/C:/Users/ydy/Desktop/llm-in-text/src/utils/convert.js#L22)
|
|
||||||
|
|
||||||
现在的问题是:
|
|
||||||
|
|
||||||
- 直接 base64 解码
|
|
||||||
- 没有严格文件大小限制
|
|
||||||
- 没有严格文件类型白名单
|
|
||||||
- 没有魔数校验
|
|
||||||
- 没有超时和并发保护
|
|
||||||
|
|
||||||
### 错误原因
|
|
||||||
|
|
||||||
当前实现是功能优先,默认相信前端传来的内容。
|
|
||||||
但上传链路是最容易出问题的地方之一。
|
|
||||||
|
|
||||||
### 会导致什么
|
|
||||||
|
|
||||||
- 大文件压垮内存
|
|
||||||
- 恶意文件拖慢 CPU
|
|
||||||
- 第三方库处理异常文件时出故障
|
|
||||||
|
|
||||||
### 整改方式
|
|
||||||
|
|
||||||
#### 对 `ocr`
|
|
||||||
|
|
||||||
1. 图片大小上限先改为 5MB 或 10MB
|
|
||||||
2. 只允许 `jpg/png/webp`
|
|
||||||
3. 服务端校验 MIME 和魔数
|
|
||||||
4. 增加请求超时
|
|
||||||
5. 增加并发限制
|
|
||||||
|
|
||||||
#### 对 `convert`
|
|
||||||
|
|
||||||
1. 只允许明确白名单格式
|
|
||||||
2. 每种格式单独设大小上限
|
|
||||||
3. 服务端检查扩展名和文件头
|
|
||||||
4. `markitdown` 执行增加超时
|
|
||||||
5. 临时文件放到独立目录
|
|
||||||
6. 临时文件异常时也要清理
|
|
||||||
|
|
||||||
### 建议白名单
|
|
||||||
|
|
||||||
- `.pdf`
|
|
||||||
- `.docx`
|
|
||||||
- `.pptx`
|
|
||||||
- `.xlsx`
|
|
||||||
- `.md`
|
|
||||||
- `.txt`
|
|
||||||
|
|
||||||
### 建议直接拒绝
|
|
||||||
|
|
||||||
- 可执行文件
|
|
||||||
- 压缩包
|
|
||||||
- 未知二进制
|
|
||||||
- 超大图片
|
|
||||||
|
|
||||||
### 验收标准
|
|
||||||
|
|
||||||
- 超限文件返回 413
|
|
||||||
- 非法类型返回 415
|
|
||||||
- OCR/convert 高并发下不会拖垮服务
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 6. 没有限流,匿名站点很容易被刷
|
|
||||||
|
|
||||||
### 具体问题
|
|
||||||
|
|
||||||
- 当前代码里没有 rate limit
|
|
||||||
- 没有按 IP、UA、路径、时间窗做限制
|
|
||||||
- `ACTIVE_COMPLETIONS` 只处理取消,不是限流器。[backend/main.py:29](/C:/Users/ydy/Desktop/llm-in-text/backend/main.py#L29)
|
|
||||||
|
|
||||||
### 错误原因
|
|
||||||
|
|
||||||
因为现在的代码默认是“正常用户正常使用”。
|
|
||||||
但匿名公网站点上线后,必须假设会被脚本反复调用。
|
|
||||||
|
|
||||||
### 会导致什么
|
|
||||||
|
|
||||||
- 模型成本失控
|
|
||||||
- CPU、内存、连接数被耗尽
|
|
||||||
- 服务变慢甚至不可用
|
|
||||||
|
|
||||||
### 整改方式
|
|
||||||
|
|
||||||
匿名站点不做登录,标准保护方式是限流:
|
|
||||||
|
|
||||||
1. 反向代理层限流
|
|
||||||
2. 应用层限流
|
|
||||||
3. 高成本接口单独限流
|
|
||||||
|
|
||||||
### 建议策略
|
|
||||||
|
|
||||||
#### `/v1/completions`
|
|
||||||
|
|
||||||
- 单 IP 每分钟 20 到 60 次
|
|
||||||
- 同时进行中的请求数限制 2 到 4 个
|
|
||||||
|
|
||||||
#### `/v1/ocr`
|
|
||||||
|
|
||||||
- 单 IP 每分钟 5 到 10 次
|
|
||||||
- 同时进行中的 OCR 限制更低
|
|
||||||
|
|
||||||
#### `/v1/convert`
|
|
||||||
|
|
||||||
- 单 IP 每分钟 3 到 5 次
|
|
||||||
- 强并发限制 1 到 2
|
|
||||||
|
|
||||||
### 还可以加什么
|
|
||||||
|
|
||||||
- Cloudflare Turnstile / hCaptcha 这类站点级防刷
|
|
||||||
- 对明显机器人流量加 challenge
|
|
||||||
|
|
||||||
这不需要登录,也不需要用户体系。
|
|
||||||
|
|
||||||
### 验收标准
|
|
||||||
|
|
||||||
- 连续脚本请求会命中 429
|
|
||||||
- 单个来源无法无限刷接口
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 7. 模型调用没有明确超时与失败策略
|
|
||||||
|
|
||||||
### 具体问题
|
|
||||||
|
|
||||||
- [backend/llm.py:15](/C:/Users/ydy/Desktop/llm-in-text/backend/llm.py#L15)
|
|
||||||
|
|
||||||
当前 `ollama.AsyncClient` 调用没有明显的统一超时和失败分类。
|
|
||||||
|
|
||||||
### 错误原因
|
|
||||||
|
|
||||||
开发阶段一般默认上游会正常返回。
|
|
||||||
但生产里,上游模型服务经常会出现:
|
|
||||||
|
|
||||||
- 变慢
|
|
||||||
- 卡住
|
|
||||||
- 连接失败
|
|
||||||
- 超时
|
|
||||||
|
|
||||||
### 会导致什么
|
|
||||||
|
|
||||||
- 请求挂很久
|
|
||||||
- 连接被占住
|
|
||||||
- 用户看起来像页面没响应
|
|
||||||
|
|
||||||
### 整改方式
|
|
||||||
|
|
||||||
1. completions 设置明确超时,例如 15 到 30 秒
|
|
||||||
2. OCR 设置更短或更明确的处理时限
|
|
||||||
3. convert 设置文件转换超时
|
|
||||||
4. 把错误分成:
|
|
||||||
- timeout
|
|
||||||
- unavailable
|
|
||||||
- bad response
|
|
||||||
5. 前端对这些错误做不同提示
|
|
||||||
|
|
||||||
### 验收标准
|
|
||||||
|
|
||||||
- 上游模型挂掉时,接口会快速失败而不是一直卡住
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 8. 缺少健康检查接口
|
|
||||||
|
|
||||||
### 具体问题
|
|
||||||
|
|
||||||
当前没有明确的:
|
|
||||||
|
|
||||||
- `/health/live`
|
|
||||||
- `/health/ready`
|
|
||||||
|
|
||||||
### 错误原因
|
|
||||||
|
|
||||||
项目还是开发态,没有进入正式部署思路。
|
|
||||||
|
|
||||||
### 会导致什么
|
|
||||||
|
|
||||||
- 你很难判断服务是否真的可用
|
|
||||||
- 容器/进程平台无法正确探活
|
|
||||||
|
|
||||||
### 整改方式
|
|
||||||
|
|
||||||
增加两个接口:
|
|
||||||
|
|
||||||
#### `/health/live`
|
|
||||||
|
|
||||||
只表示“应用进程活着”
|
|
||||||
|
|
||||||
#### `/health/ready`
|
|
||||||
|
|
||||||
表示“应用准备好服务请求”
|
|
||||||
|
|
||||||
这个接口至少检查:
|
|
||||||
|
|
||||||
- 模型上游是否可连接
|
|
||||||
- 关键配置是否存在
|
|
||||||
|
|
||||||
### 验收标准
|
|
||||||
|
|
||||||
- 反向代理或容器平台能用它判断是否接流量
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 9. 预览组件有 XSS 风险
|
|
||||||
|
|
||||||
### 具体问题
|
|
||||||
|
|
||||||
- [src/components/MarkdownPreview.vue:2](/C:/Users/ydy/Desktop/llm-in-text/src/components/MarkdownPreview.vue#L2)
|
|
||||||
- [src/components/MarkdownPreview.vue:19](/C:/Users/ydy/Desktop/llm-in-text/src/components/MarkdownPreview.vue#L19)
|
|
||||||
|
|
||||||
现在做法是:
|
|
||||||
|
|
||||||
- `v-html`
|
|
||||||
- `html: true`
|
|
||||||
|
|
||||||
### 错误原因
|
|
||||||
|
|
||||||
这意味着 markdown 里的原始 HTML 会被直接渲染。
|
|
||||||
如果内容来源不完全可信,这就是典型 XSS 入口。
|
|
||||||
|
|
||||||
### 会导致什么
|
|
||||||
|
|
||||||
- 恶意脚本执行
|
|
||||||
- 页面被注入恶意 DOM
|
|
||||||
|
|
||||||
### 整改方式
|
|
||||||
|
|
||||||
你有两个选择:
|
|
||||||
|
|
||||||
#### 方案 A:最简单,直接关掉 HTML
|
|
||||||
|
|
||||||
把:
|
|
||||||
|
|
||||||
```js
|
|
||||||
html: true
|
|
||||||
```
|
|
||||||
|
|
||||||
改成:
|
|
||||||
|
|
||||||
```js
|
|
||||||
html: false
|
|
||||||
```
|
|
||||||
|
|
||||||
#### 方案 B:保留 HTML,但做净化
|
|
||||||
|
|
||||||
1. 引入 DOMPurify
|
|
||||||
2. `md.render()` 后先 sanitize
|
|
||||||
3. 再给 `v-html`
|
|
||||||
|
|
||||||
### 推荐
|
|
||||||
|
|
||||||
如果你不是必须支持原始 HTML,直接用方案 A。
|
|
||||||
|
|
||||||
### 验收标准
|
|
||||||
|
|
||||||
- 恶意 markdown/HTML 不会执行脚本
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 10. 前端默认配置会误连固定服务地址
|
|
||||||
|
|
||||||
### 具体问题
|
|
||||||
|
|
||||||
- [src/utils/config.js:1](/C:/Users/ydy/Desktop/llm-in-text/src/utils/config.js#L1)
|
|
||||||
- [.env.example:1](/C:/Users/ydy/Desktop/llm-in-text/.env.example#L1)
|
|
||||||
- [backend/llm.py:12](/C:/Users/ydy/Desktop/llm-in-text/backend/llm.py#L12)
|
|
||||||
|
|
||||||
默认值里有固定公网域名和固定内网 IP。
|
|
||||||
|
|
||||||
### 错误原因
|
|
||||||
|
|
||||||
这是把“某次部署环境”写成了“代码默认值”。
|
|
||||||
|
|
||||||
### 会导致什么
|
|
||||||
|
|
||||||
- 本地开发可能误连生产或旧环境
|
|
||||||
- 新服务器部署时容易配错
|
|
||||||
|
|
||||||
### 整改方式
|
|
||||||
|
|
||||||
1. 默认值改成本地开发地址或空值
|
|
||||||
2. 关键配置不存在时直接报错
|
|
||||||
3. `.env.example` 只放模板,不放真实地址
|
|
||||||
|
|
||||||
### 建议
|
|
||||||
|
|
||||||
- `VITE_API_BASE_URL` 默认走同域,如 `''`
|
|
||||||
- 前端优先使用 `/v1/...` 反代
|
|
||||||
- 后端 `OLLAMA_HOST` 必须来自环境变量
|
|
||||||
|
|
||||||
### 验收标准
|
|
||||||
|
|
||||||
- 不配置环境变量时,不会误连旧服务
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 11. 包版本和界面版本不一致
|
|
||||||
|
|
||||||
### 具体问题
|
|
||||||
|
|
||||||
- [package.json:4](/C:/Users/ydy/Desktop/llm-in-text/package.json#L4) 是 `0.0.0`
|
|
||||||
- [src/components/SettingsPanel.vue:275](/C:/Users/ydy/Desktop/llm-in-text/src/components/SettingsPanel.vue#L275) 写的是 `v0.1.0-beta`
|
|
||||||
|
|
||||||
### 错误原因
|
|
||||||
|
|
||||||
一个是包元数据,一个是手写展示文案,没人保证同步。
|
|
||||||
|
|
||||||
### 会导致什么
|
|
||||||
|
|
||||||
- 发布后你都不确定线上到底是哪版
|
|
||||||
|
|
||||||
### 整改方式
|
|
||||||
|
|
||||||
1. 统一从 `package.json` 注入版本
|
|
||||||
2. 前端不要手写版本号
|
|
||||||
|
|
||||||
### 验收标准
|
|
||||||
|
|
||||||
- 页面显示版本和构建版本完全一致
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 12. `package.json` 缺少质量脚本
|
|
||||||
|
|
||||||
### 具体问题
|
|
||||||
|
|
||||||
- [package.json:6](/C:/Users/ydy/Desktop/llm-in-text/package.json#L6)
|
|
||||||
|
|
||||||
当前只有:
|
|
||||||
|
|
||||||
- `dev`
|
|
||||||
- `build`
|
|
||||||
- `preview`
|
|
||||||
|
|
||||||
没有:
|
|
||||||
|
|
||||||
- `test`
|
|
||||||
- `lint`
|
|
||||||
- `check`
|
|
||||||
|
|
||||||
### 错误原因
|
|
||||||
|
|
||||||
项目还停留在“能运行”的阶段,没有建立质量门禁。
|
|
||||||
|
|
||||||
### 会导致什么
|
|
||||||
|
|
||||||
- 任何改动都只能靠手工试
|
|
||||||
- 回归问题容易漏
|
|
||||||
|
|
||||||
### 整改方式
|
|
||||||
|
|
||||||
至少补这些脚本:
|
|
||||||
|
|
||||||
```json
|
|
||||||
"scripts": {
|
|
||||||
"dev": "vite",
|
|
||||||
"build": "vite build",
|
|
||||||
"preview": "vite preview",
|
|
||||||
"test": "pytest backend/tests -q",
|
|
||||||
"lint": "eslint src",
|
|
||||||
"check": "npm run lint && npm run build && pytest backend/tests -q"
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
如果前端暂时没配 ESLint,也至少先把 `test` 和 `check` 建起来。
|
|
||||||
|
|
||||||
### 验收标准
|
|
||||||
|
|
||||||
- 以后每次改代码前后都能统一执行 `check`
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 13. 测试覆盖不够,缺关键路径
|
|
||||||
|
|
||||||
### 具体问题
|
|
||||||
|
|
||||||
当前测试主要是:
|
|
||||||
|
|
||||||
- prompt 构造
|
|
||||||
- 取消逻辑
|
|
||||||
- LLM 消息结构
|
|
||||||
|
|
||||||
缺少:
|
|
||||||
|
|
||||||
- 匿名访问基本流程
|
|
||||||
- 错误响应格式
|
|
||||||
- OCR 文件限制
|
|
||||||
- convert 文件限制
|
|
||||||
- 限流
|
|
||||||
- XSS
|
|
||||||
|
|
||||||
### 错误原因
|
|
||||||
|
|
||||||
现有测试更偏功能开发时的局部验证,不是上线前测试矩阵。
|
|
||||||
|
|
||||||
### 整改方式
|
|
||||||
|
|
||||||
补这些测试:
|
|
||||||
|
|
||||||
1. completions 正常返回
|
|
||||||
2. completions 上游超时
|
|
||||||
3. completions 请求超长
|
|
||||||
4. OCR 非法类型
|
|
||||||
5. OCR 超大图片
|
|
||||||
6. convert 非法类型
|
|
||||||
7. convert 超大文件
|
|
||||||
8. 限流命中
|
|
||||||
9. 未授权方案移除后,匿名访问可正常工作
|
|
||||||
10. markdown 预览 XSS 样例
|
|
||||||
|
|
||||||
### 验收标准
|
|
||||||
|
|
||||||
- 关键错误分支都有自动化测试
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 14. Service Worker 缓存策略还不够稳
|
|
||||||
|
|
||||||
### 具体问题
|
|
||||||
|
|
||||||
- [src/main.js:13](/C:/Users/ydy/Desktop/llm-in-text/src/main.js#L13)
|
|
||||||
- [public/sw.js:1](/C:/Users/ydy/Desktop/llm-in-text/public/sw.js#L1)
|
|
||||||
|
|
||||||
当前是手写缓存逻辑,版本固定写死。
|
|
||||||
|
|
||||||
### 错误原因
|
|
||||||
|
|
||||||
这是一个能用的基础实现,但不适合长期生产维护。
|
|
||||||
|
|
||||||
### 会导致什么
|
|
||||||
|
|
||||||
- 更新后可能缓存混乱
|
|
||||||
- 老版本资源残留
|
|
||||||
|
|
||||||
### 整改方式
|
|
||||||
|
|
||||||
如果你不强依赖离线能力:
|
|
||||||
|
|
||||||
1. 先临时关闭 SW
|
|
||||||
2. 等核心功能稳定后再重做 PWA
|
|
||||||
|
|
||||||
如果要保留:
|
|
||||||
|
|
||||||
1. 用成熟方案接管,如 Vite PWA / Workbox
|
|
||||||
2. 资源按 hash 控制
|
|
||||||
3. 做更新提示
|
|
||||||
|
|
||||||
### 推荐
|
|
||||||
|
|
||||||
如果现在重点是先上线稳定版,先停掉 SW 更省事。
|
|
||||||
|
|
||||||
### 验收标准
|
|
||||||
|
|
||||||
- 用户刷新后不会出现随机旧资源
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 15. 构建体积偏大
|
|
||||||
|
|
||||||
### 具体问题
|
|
||||||
|
|
||||||
本次构建已经出现大 chunk 警告,尤其是 Mermaid 相关包比较重。
|
|
||||||
|
|
||||||
### 错误原因
|
|
||||||
|
|
||||||
图表、编辑器、语法高亮、数学渲染这类库本身就大。
|
|
||||||
现在又没有足够按功能懒加载。
|
|
||||||
|
|
||||||
### 会导致什么
|
|
||||||
|
|
||||||
- 首屏慢
|
|
||||||
- 弱网体验差
|
|
||||||
|
|
||||||
### 整改方式
|
|
||||||
|
|
||||||
1. Mermaid 按需加载
|
|
||||||
2. 预览按需加载
|
|
||||||
3. OCR / convert 相关 UI 按需加载
|
|
||||||
4. 收敛 `manualChunks`
|
|
||||||
|
|
||||||
### 验收标准
|
|
||||||
|
|
||||||
- 首页首次加载明显更轻
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 16. 匿名站点应该怎么做保护,而不是登录
|
|
||||||
|
|
||||||
这是你这个项目最关键的方向问题。
|
|
||||||
|
|
||||||
你不想做用户级网站,这完全可以。
|
|
||||||
|
|
||||||
那就按匿名站点的标准做:
|
|
||||||
|
|
||||||
### 必做
|
|
||||||
|
|
||||||
1. 去掉前端共享密钥
|
|
||||||
2. 收紧 CORS
|
|
||||||
3. 加 Nginx / Cloudflare / 网关限流
|
|
||||||
4. 应用层再做限流
|
|
||||||
5. 限制请求体大小
|
|
||||||
6. 限制 OCR/convert 并发
|
|
||||||
7. 错误信息脱敏
|
|
||||||
8. 加健康检查
|
|
||||||
9. 处理 XSS
|
|
||||||
|
|
||||||
### 可选
|
|
||||||
|
|
||||||
1. Cloudflare Turnstile
|
|
||||||
2. 简单的人机验证 challenge
|
|
||||||
3. 对高频匿名流量启用冷却时间
|
|
||||||
|
|
||||||
### 不必做
|
|
||||||
|
|
||||||
1. 登录
|
|
||||||
2. 注册
|
|
||||||
3. 用户系统
|
|
||||||
4. JWT
|
|
||||||
|
|
||||||
只要你的目标是匿名工具站,而不是多租户平台,上面这套就够了。
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 最简修复顺序
|
|
||||||
|
|
||||||
如果你要最低成本把项目拉到“能较安全公开上线”的程度,建议顺序是:
|
|
||||||
|
|
||||||
1. 删除前端 API key 和后端固定 key 校验
|
|
||||||
2. 删除 IP 获取和地理位置推断
|
|
||||||
3. 收紧 CORS
|
|
||||||
4. 统一错误响应
|
|
||||||
5. 给 OCR/convert 加大小、类型、超时限制
|
|
||||||
6. 加限流
|
|
||||||
7. 修掉 `MarkdownPreview` 的 XSS 风险
|
|
||||||
8. 增加 `/health/live` 和 `/health/ready`
|
|
||||||
9. 补 `test` / `check` 脚本
|
|
||||||
10. 视情况先关闭 service worker
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 这份清单对应的文件
|
|
||||||
|
|
||||||
- [backend/main.py](/C:/Users/ydy/Desktop/llm-in-text/backend/main.py)
|
|
||||||
- [backend/llm.py](/C:/Users/ydy/Desktop/llm-in-text/backend/llm.py)
|
|
||||||
- [src/utils/api.js](/C:/Users/ydy/Desktop/llm-in-text/src/utils/api.js)
|
|
||||||
- [src/utils/convert.js](/C:/Users/ydy/Desktop/llm-in-text/src/utils/convert.js)
|
|
||||||
- [src/components/MarkdownPreview.vue](/C:/Users/ydy/Desktop/llm-in-text/src/components/MarkdownPreview.vue)
|
|
||||||
- [src/stores/settings.js](/C:/Users/ydy/Desktop/llm-in-text/src/stores/settings.js)
|
|
||||||
- [src/components/SettingsPanel.vue](/C:/Users/ydy/Desktop/llm-in-text/src/components/SettingsPanel.vue)
|
|
||||||
- [public/sw.js](/C:/Users/ydy/Desktop/llm-in-text/public/sw.js)
|
|
||||||
- [package.json](/C:/Users/ydy/Desktop/llm-in-text/package.json)
|
|
||||||
|
|
||||||
@@ -1,579 +0,0 @@
|
|||||||
# llm-in-text 生产环境修复清单
|
|
||||||
|
|
||||||
## 文档目的
|
|
||||||
|
|
||||||
本文档是针对当前仓库在 2026-04-01 状态下基于代码的生产就绪性审查。
|
|
||||||
|
|
||||||
它回答三个问题:
|
|
||||||
|
|
||||||
1. 当前项目是否已准备好投入生产?
|
|
||||||
2. 自上次审查以来已修复了哪些问题?
|
|
||||||
3. 还有什么因素阻碍安全上线?
|
|
||||||
|
|
||||||
本次审查仅限于仓库中已有的内容和本地直接验证的内容,不包括外部基础设施、反向代理配置、云资源、CI 平台密钥或运行时运维的完整审计。
|
|
||||||
|
|
||||||
## 审查范围
|
|
||||||
|
|
||||||
- 前端构建和运行时入口
|
|
||||||
- 后端 FastAPI 端点和模型集成
|
|
||||||
- 补全、OCR 和文件转换请求路径
|
|
||||||
- 隐私相关行为和本地存储
|
|
||||||
- 基础测试覆盖率和构建验证
|
|
||||||
- PWA/Service Worker 状态
|
|
||||||
- 仓库中存在的部署和运维工件
|
|
||||||
|
|
||||||
## 已验证的事实
|
|
||||||
|
|
||||||
以下项目已针对当前仓库直接验证:
|
|
||||||
|
|
||||||
- 前端生产构建通过 `npm.cmd run build` 成功
|
|
||||||
- 后端测试通过 `pytest backend/tests -q`
|
|
||||||
- 当前后端测试结果:`8 passed, 1 skipped`
|
|
||||||
- 存在健康检查端点:`/health/live` 和 `/health/ready`
|
|
||||||
- 前端不再硬编码 API 密钥
|
|
||||||
- 后端不再需要旧的 `X-API-Key` 头
|
|
||||||
- 前端隐私模式现在默认为启用
|
|
||||||
- 后端错误响应现在已规范化,不再返回原始异常字符串
|
|
||||||
- OCR 和转换端点现在强制执行基本的文件大小和扩展名检查
|
|
||||||
|
|
||||||
## 当前评估
|
|
||||||
|
|
||||||
项目尚未准备好投入生产。
|
|
||||||
|
|
||||||
当前状态更接近于:
|
|
||||||
|
|
||||||
- 一个可用的原型
|
|
||||||
- 内部演示
|
|
||||||
- 具有部分加固的预发布候选版本
|
|
||||||
|
|
||||||
由于缺少几个核心生产基线,它尚未准备好面向互联网的生产使用:
|
|
||||||
|
|
||||||
- 真正的限流和并发保护
|
|
||||||
- 正式的部署工件和运行时拓扑
|
|
||||||
- CI/CD 和自动化质量门禁
|
|
||||||
- 结构化的可观测性和告警
|
|
||||||
- 更强的请求验证和更安全的文件处理隔离
|
|
||||||
- 前端自动化测试和端到端验证
|
|
||||||
|
|
||||||
## 已修复的问题
|
|
||||||
|
|
||||||
与之前的清单相比,以下项目不再作为阻碍因素:
|
|
||||||
|
|
||||||
### 已修复:硬编码的前端/后端共享 API 密钥
|
|
||||||
|
|
||||||
- `src/utils/api.js` 不再发送 `X-API-Key`
|
|
||||||
- `src/utils/convert.js` 不再发送 `X-API-Key`
|
|
||||||
- `backend/main.py` 不再强制执行旧的静态密钥
|
|
||||||
|
|
||||||
这消除了上次审查中最严重的问题之一。
|
|
||||||
|
|
||||||
遗留问题:
|
|
||||||
|
|
||||||
- `backend/tests/test_main_cancel.py` 仍然包含过时的 `X-API-Key` 头,但它们目前是无效的,表明测试假设已过时而非活动中的认证逻辑
|
|
||||||
|
|
||||||
### 已修复:危险的通配符 CORS 配置
|
|
||||||
|
|
||||||
当前后端 CORS 限制为:
|
|
||||||
|
|
||||||
- `http://localhost:5173`
|
|
||||||
- `http://localhost:3000`
|
|
||||||
|
|
||||||
并使用:
|
|
||||||
|
|
||||||
- `allow_credentials=False`
|
|
||||||
- `allow_methods=["POST", "OPTIONS"]`
|
|
||||||
- `allow_headers=["Content-Type", "X-Request-Id"]`
|
|
||||||
|
|
||||||
这比之前的通配符配置安全得多。
|
|
||||||
|
|
||||||
遗留问题:
|
|
||||||
|
|
||||||
- CORS 仍然针对本地开发硬编码,对于 staging/生产环境不是环境驱动的
|
|
||||||
|
|
||||||
### 已修复:默认收集前端公网 IP
|
|
||||||
|
|
||||||
- `src/stores/settings.js` 现在将 `privacyMode` 默认为 `true`
|
|
||||||
- `src/utils/api.js` 不再调用 `ipify`
|
|
||||||
- `src/utils/api.js` 不再发送 `X-Client-IP`
|
|
||||||
- `backend/main.py` 不再将 IP 派生的位置注入提示词
|
|
||||||
|
|
||||||
遗留问题:
|
|
||||||
|
|
||||||
- `backend/geoip.py` 和 GeoLite 数据库仍然存在于仓库中,这可能会造成对当前隐私模型的混淆
|
|
||||||
|
|
||||||
### 已修复:向客户端暴露原始异常字符串
|
|
||||||
|
|
||||||
当前后端响应使用 `_error_response(...)` 和结构化载荷,例如:
|
|
||||||
|
|
||||||
- `INTERNAL_ERROR`
|
|
||||||
- `OCR_FAILED`
|
|
||||||
- `CONVERT_FAILED`
|
|
||||||
- `FILE_TOO_LARGE`
|
|
||||||
- `INVALID_FILE_TYPE`
|
|
||||||
|
|
||||||
这比直接返回 `str(exception)` 更好。
|
|
||||||
|
|
||||||
遗留问题:
|
|
||||||
|
|
||||||
- 日志仍然记录用户派生内容的预览,这是另一个隐私/可观测性问题
|
|
||||||
|
|
||||||
### 部分修复:上传和转换输入边界
|
|
||||||
|
|
||||||
`backend/main.py` 中的当前后端保护包括:
|
|
||||||
|
|
||||||
- OCR 大小上限:10 MB
|
|
||||||
- 转换大小上限:50 MB
|
|
||||||
- OCR 和转换的扩展名白名单
|
|
||||||
- 在 `finally` 中清理临时转换文件
|
|
||||||
|
|
||||||
这是有意义的进展,但不足以投入生产。
|
|
||||||
|
|
||||||
## 生产阻碍因素
|
|
||||||
|
|
||||||
优先级含义:
|
|
||||||
|
|
||||||
- `P0`:必须在生产上线前完成
|
|
||||||
- `P1`:应在公开发布或广泛推广前完成
|
|
||||||
- `P2`:重要的后续加固和维护工作
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## P0 阻碍因素
|
|
||||||
|
|
||||||
### P0-01 声明了限流和并发控制但未强制执行
|
|
||||||
|
|
||||||
当前状态:
|
|
||||||
|
|
||||||
- `backend/main.py` 定义了 `MAX_CONCURRENT_COMPLETIONS = 4`
|
|
||||||
- `backend/main.py` 定义了 `COMPLETION_RATE_LIMIT = 60`
|
|
||||||
- 没有任何实际的限流器使用这两个值
|
|
||||||
- OCR 和转换端点也没有真正的每客户端节流
|
|
||||||
|
|
||||||
风险:
|
|
||||||
|
|
||||||
- 容易滥用昂贵的补全/OCR/转换端点
|
|
||||||
- 可避免地对模型主机、CPU、内存和临时存储造成过载
|
|
||||||
- 突发流量下无法控制降级
|
|
||||||
|
|
||||||
必需的修复:
|
|
||||||
|
|
||||||
1. 在应用或网关中强制执行真正的每路由限流
|
|
||||||
2. 为补全作业添加真正的并发保护
|
|
||||||
3. 为 `/v1/completions`、`/v1/ocr`、`/v1/convert` 添加单独预算
|
|
||||||
4. 达到限制时返回明确的 `429`
|
|
||||||
5. 为节流的请求和队列深度发出指标
|
|
||||||
|
|
||||||
验收标准:
|
|
||||||
|
|
||||||
- 重复的突发流量触发确定性的 `429`
|
|
||||||
- 并发补全不能超过配置的预算
|
|
||||||
- 负载测试下服务保持稳定
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
### P0-02 仓库中不存在生产部署基线
|
|
||||||
|
|
||||||
当前状态:
|
|
||||||
|
|
||||||
- 后端仅通过 `uvicorn.run(...)` 暴露开发式启动
|
|
||||||
- 没有 `Dockerfile`
|
|
||||||
- 没有 compose 文件
|
|
||||||
- 没有 Kubernetes 清单或 Helm chart
|
|
||||||
- 没有 systemd 单元
|
|
||||||
- 没有反向代理参考配置
|
|
||||||
- 没有记录的生产环境契约
|
|
||||||
|
|
||||||
风险:
|
|
||||||
|
|
||||||
- 没有可复现的部署路径
|
|
||||||
- 没有明确的过程监督、重启或优雅的发布模型
|
|
||||||
- 没有记录的 ingress/请求体大小/超时/TLS 姿态
|
|
||||||
|
|
||||||
必需的修复:
|
|
||||||
|
|
||||||
1. 定义一个官方部署目标
|
|
||||||
2. 为该目标添加部署工件
|
|
||||||
3. 记录所需的环境变量、端口、探针和存储
|
|
||||||
4. 定义优雅关闭和发布行为
|
|
||||||
5. 记录反向代理限制和信任边界
|
|
||||||
|
|
||||||
验收标准:
|
|
||||||
|
|
||||||
- 新环境可以从仓库文档和工件部署
|
|
||||||
- 健康探针已接入所选运行时
|
|
||||||
- 回滚路径已记录
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
### P0-03 缺少 CI/CD 和仓库质量门禁
|
|
||||||
|
|
||||||
当前状态:
|
|
||||||
|
|
||||||
- 仓库级别没有找到 `.github/workflows`
|
|
||||||
- `package.json` 没有真正的前端测试脚本
|
|
||||||
- 没有 lint 脚本
|
|
||||||
- 没有类型检查脚本
|
|
||||||
- 没有依赖扫描或密钥扫描工作流
|
|
||||||
|
|
||||||
风险:
|
|
||||||
|
|
||||||
- 回归只能手动捕获
|
|
||||||
- 安全和打包漂移很可能发生
|
|
||||||
- 生产就绪性取决于本地开发者的规范
|
|
||||||
|
|
||||||
必需的修复:
|
|
||||||
|
|
||||||
1. 为构建和测试添加 CI 工作流
|
|
||||||
2. 添加前端自动化测试
|
|
||||||
3. 在适用的情况下添加 lint 和类型检查门禁
|
|
||||||
4. 添加依赖漏洞扫描
|
|
||||||
5. 添加密钥扫描和基本 SAST
|
|
||||||
|
|
||||||
验收标准:
|
|
||||||
|
|
||||||
- 每个 PR 都运行构建、后端测试、前端测试和 lint
|
|
||||||
- 失败的检查阻止合并
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
### P0-04 文件处理路径仍然缺乏生产级隔离
|
|
||||||
|
|
||||||
当前状态:
|
|
||||||
|
|
||||||
- 后端将完整 base64 载荷解码到内存中
|
|
||||||
- 转换写入临时文件并将其传递给 `markitdown`
|
|
||||||
- OCR 和转换主要依赖扩展名检查,而不是内容嗅探
|
|
||||||
- 没有工作进程隔离或用于转换的单独沙箱
|
|
||||||
- 转换任务没有队列或资源预算
|
|
||||||
|
|
||||||
风险:
|
|
||||||
|
|
||||||
- 大型或并发上传导致内存峰值
|
|
||||||
- 畸形或对抗性文档会给解析器带来压力
|
|
||||||
- 转换工作负载可能干扰核心补全可用性
|
|
||||||
|
|
||||||
必需的修复:
|
|
||||||
|
|
||||||
1. 明确验证 base64 解码失败
|
|
||||||
2. 添加 MIME/内容嗅探,而不仅仅是扩展名检查
|
|
||||||
3. 在入口和应用层添加更低的、路由特定的请求体限制
|
|
||||||
4. 将转换隔离到单独的工作进程/进程边界
|
|
||||||
5. 添加超时、并发上限和队列深度控制
|
|
||||||
|
|
||||||
验收标准:
|
|
||||||
|
|
||||||
- 畸形载荷失败并返回明确的 4xx 响应
|
|
||||||
- 转换不能饿死补全服务
|
|
||||||
- 压力下临时文件和内存增长保持有界
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
### P0-05 环境配置不一致且不安全
|
|
||||||
|
|
||||||
当前状态:
|
|
||||||
|
|
||||||
- 前端 `.env.example` 相当安全
|
|
||||||
- 后端 `.env.example` 过时且与代码不一致
|
|
||||||
- 代码读取 `OLLAMA_HOST`
|
|
||||||
- 后端示例仍然使用 `OLLAMA_BASE_URL`
|
|
||||||
- 后端示例仍然定义 `OPENAI_API_KEY=ollama`,这是误导性的
|
|
||||||
|
|
||||||
风险:
|
|
||||||
|
|
||||||
- 新环境配置不正确
|
|
||||||
- 运营商可能假设不支持的认证/配置行为
|
|
||||||
- staging/生产漂移很可能发生
|
|
||||||
|
|
||||||
必需的修复:
|
|
||||||
|
|
||||||
1. 用代码实际使用的变量替换后端环境示例
|
|
||||||
2. 在启动时验证所需的环境变量
|
|
||||||
3. 分离开发、staging 和生产环境契约
|
|
||||||
4. 缺失关键配置时快速失败
|
|
||||||
|
|
||||||
验收标准:
|
|
||||||
|
|
||||||
- 示例环境文件匹配真实运行时行为
|
|
||||||
- 无效或缺失关键配置导致启动失败
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## P1 高优先级差距
|
|
||||||
|
|
||||||
### P1-01 日志仍然捕获用户派生内容预览
|
|
||||||
|
|
||||||
当前状态:
|
|
||||||
|
|
||||||
- `backend/main.py` 记录提示词派生的前缀和后缀预览
|
|
||||||
- 补全结果记录包含内容预览
|
|
||||||
- OCR 和转换记录文本预览长度和片段
|
|
||||||
|
|
||||||
风险:
|
|
||||||
|
|
||||||
- 日志可能包含敏感文档内容
|
|
||||||
- 隐私姿态与应用可见的隐私设置不一致
|
|
||||||
- 难以证明保留/合规姿态
|
|
||||||
|
|
||||||
必需的修复:
|
|
||||||
|
|
||||||
1. 默认情况下停止记录用户内容正文和预览
|
|
||||||
2. 仅保留请求元数据:路由、请求 ID、状态、延迟、大小
|
|
||||||
3. 引入结构化 JSON 日志
|
|
||||||
4. 脱敏或哈希任何敏感标识符
|
|
||||||
|
|
||||||
验收标准:
|
|
||||||
|
|
||||||
- 默认日志不包含用户文档文本
|
|
||||||
- 请求关联仍可通过请求 ID 和元数据工作
|
|
||||||
|
|
||||||
### P1-02 请求验证仍然过于宽松
|
|
||||||
|
|
||||||
当前状态:
|
|
||||||
|
|
||||||
- Pydantic 模型定义了字段但没有长度或枚举约束
|
|
||||||
- `prefix`、`suffix`、`filename` 和 `reason` 受到极小约束
|
|
||||||
- base64 字段在路由特定检查之前仍然可能非常大
|
|
||||||
- 前端转换路径在将完整文件读入 base64 之前不执行预验证
|
|
||||||
|
|
||||||
风险:
|
|
||||||
|
|
||||||
- 过大或畸形的请求太容易到达昂贵的逻辑
|
|
||||||
- 端点之间的 4xx 行为不一致
|
|
||||||
|
|
||||||
必需的修复:
|
|
||||||
|
|
||||||
1. 为 Pydantic 字段添加长度和枚举约束
|
|
||||||
2. 明确验证 base64 格式
|
|
||||||
3. 更防御性地规范化文件名处理
|
|
||||||
4. 添加前端预检查大小/类型作为 UX
|
|
||||||
|
|
||||||
验收标准:
|
|
||||||
|
|
||||||
- 畸形请求尽早失败并返回确定性的 4xx 响应
|
|
||||||
|
|
||||||
### P1-03 前端自动化覆盖率基本缺失
|
|
||||||
|
|
||||||
当前状态:
|
|
||||||
|
|
||||||
- 后端有针对性的单元/集成风格测试
|
|
||||||
- 前端没有配置测试运行器
|
|
||||||
- 核心用户路径没有 E2E 覆盖率
|
|
||||||
|
|
||||||
重要说明:
|
|
||||||
|
|
||||||
- `backend/tests/test_main_cancel.py` 仍然发送过时的 `X-API-Key` 头;测试通过仅仅是因为后端忽略它们
|
|
||||||
|
|
||||||
风险:
|
|
||||||
|
|
||||||
- 编辑器、上传、OCR、转换和设置的回归将会遗漏
|
|
||||||
|
|
||||||
必需的修复:
|
|
||||||
|
|
||||||
1. 添加前端单元/组件测试
|
|
||||||
2. 为补全、取消、上传、OCR 和转换添加 E2E 覆盖率
|
|
||||||
3. 从后端测试中移除过时的认证假设
|
|
||||||
|
|
||||||
验收标准:
|
|
||||||
|
|
||||||
- 核心用户旅程在 CI 中自动覆盖
|
|
||||||
|
|
||||||
### P1-04 健康检查端点存在,但就绪性浅且缺少可观测性
|
|
||||||
|
|
||||||
当前状态:
|
|
||||||
|
|
||||||
- `/health/live` 存在
|
|
||||||
- `/health/ready` 存在
|
|
||||||
- 就绪性实际上不检查上游模型可用性
|
|
||||||
- 没有指标端点
|
|
||||||
- 没有追踪
|
|
||||||
- 没有告警定义
|
|
||||||
|
|
||||||
风险:
|
|
||||||
|
|
||||||
- 运行时故障检测太晚
|
|
||||||
- 平台探针可能报告健康而上游依赖不可用
|
|
||||||
|
|
||||||
必需的修复:
|
|
||||||
|
|
||||||
1. 使就绪性反映关键依赖状态
|
|
||||||
2. 添加请求/延迟/错误指标
|
|
||||||
3. 为上游故障和饱和添加告警阈值
|
|
||||||
4. 为关键路由定义仪表板
|
|
||||||
|
|
||||||
验收标准:
|
|
||||||
|
|
||||||
- 运营商可以快速检测模型依赖故障
|
|
||||||
- 请求成功率和延迟可观测
|
|
||||||
|
|
||||||
### P1-05 Service Worker 实现存在但被禁用
|
|
||||||
|
|
||||||
当前状态:
|
|
||||||
|
|
||||||
- `public/sw.js` 存在
|
|
||||||
- `src/main.js` 用 `&& false` 硬禁用注册
|
|
||||||
- Service Worker 策略是手写的并通过静态缓存名称版本化
|
|
||||||
|
|
||||||
风险:
|
|
||||||
|
|
||||||
- 当前仓库包含未被实际使用的休眠 PWA 逻辑
|
|
||||||
- 如果随意重新启用,更新和缓存行为可能很脆弱
|
|
||||||
|
|
||||||
必需的修复:
|
|
||||||
|
|
||||||
1. 确定 PWA 是否在生产范围内
|
|
||||||
2. 如果是,采用维护的策略如 Vite PWA/Workbox
|
|
||||||
3. 如果不是,移除无用的 Service Worker 代码以减少混淆
|
|
||||||
|
|
||||||
验收标准:
|
|
||||||
|
|
||||||
- PWA 行为要么被有意支持和测试,要么被完全移除
|
|
||||||
|
|
||||||
### P1-06 背景图像持久化可能导致本地存储和内存膨胀
|
|
||||||
|
|
||||||
当前状态:
|
|
||||||
|
|
||||||
- 设置面板将上传的背景图像读取为 data URL
|
|
||||||
- 背景图像数据存储在 localStorage 中
|
|
||||||
- 没有对背景资产强制执行明确的大小上限
|
|
||||||
|
|
||||||
风险:
|
|
||||||
|
|
||||||
- 存储配额耗尽
|
|
||||||
- 大图像导致 UI 缓慢
|
|
||||||
- 跨浏览器的持久化行为脆弱
|
|
||||||
|
|
||||||
必需的修复:
|
|
||||||
|
|
||||||
1. 读取前在客户端添加大小上限
|
|
||||||
2. 持久化前调整大小/压缩
|
|
||||||
3. 对于较大的资产,优先使用 blob/object URL 或 IndexedDB
|
|
||||||
4. 为存储溢出添加迁移/错误处理
|
|
||||||
|
|
||||||
验收标准:
|
|
||||||
|
|
||||||
- 大图像不能降低启动或破坏设置持久化
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## P2 重要后续工作
|
|
||||||
|
|
||||||
### P2-01 构建成功,但 bundle/chunk 策略仍然粗糙
|
|
||||||
|
|
||||||
验证的构建输出显示:
|
|
||||||
|
|
||||||
- `manualChunks` 生成许多空 chunk
|
|
||||||
- 一个与 Mermaid 相关的大型 chunk 超过 1 MB 压缩后
|
|
||||||
|
|
||||||
风险:
|
|
||||||
|
|
||||||
- 不必要的 chunk 开销
|
|
||||||
- 较弱的设备上冷启动较慢
|
|
||||||
|
|
||||||
必需的修复:
|
|
||||||
|
|
||||||
1. 简化 `manualChunks`
|
|
||||||
2. 延迟加载重型可选功能
|
|
||||||
3. 清理 chunk 后重新测量首次加载成本
|
|
||||||
|
|
||||||
### P2-02 OCR 缓存和图像哈希缓存没有明确的驱逐策略
|
|
||||||
|
|
||||||
当前状态:
|
|
||||||
|
|
||||||
- OCR 数据存储在内存中的 `Map`
|
|
||||||
- 哈希缓存也在内存中
|
|
||||||
- 没有 TTL
|
|
||||||
- 没有最大条目数
|
|
||||||
|
|
||||||
风险:
|
|
||||||
|
|
||||||
- 长时间会话会累积内存
|
|
||||||
|
|
||||||
必需的修复:
|
|
||||||
|
|
||||||
1. 添加 TTL 和条目边界
|
|
||||||
2. 如需要,暴露缓存指标用于调试
|
|
||||||
|
|
||||||
### P2-03 仓库仍然包含过时和混淆的工件
|
|
||||||
|
|
||||||
示例:
|
|
||||||
|
|
||||||
- `backend/geoip.py` 和 GeoLite DB 仍然存在,尽管 IP 地理定位在请求流程中不再活跃
|
|
||||||
- `backend/.env.example` 记录的变量不是代码使用的
|
|
||||||
- 后端测试仍然包含过时的 `X-API-Key`
|
|
||||||
|
|
||||||
风险:
|
|
||||||
|
|
||||||
- 未来维护者可能无意中重新引入已移除的行为
|
|
||||||
|
|
||||||
必需的修复:
|
|
||||||
|
|
||||||
1. 移除死代码和过时配置
|
|
||||||
2. 使测试和文档与当前实现保持一致
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 上线前必需的缺失证据
|
|
||||||
|
|
||||||
仓库目前不提供以下生产能力的证据:
|
|
||||||
|
|
||||||
- staging 部署管道
|
|
||||||
- 回滚程序
|
|
||||||
- 流量/负载测试结果
|
|
||||||
- 故障注入或混沌测试
|
|
||||||
- 备份/恢复程序
|
|
||||||
- 事件响应运行手册
|
|
||||||
- SLO/SLA 定义
|
|
||||||
- 安全扫描基线
|
|
||||||
- 依赖更新策略
|
|
||||||
- 隐私/数据保留文档
|
|
||||||
|
|
||||||
目前应将证据缺失视为未就绪,而不是隐式完成。
|
|
||||||
|
|
||||||
## 推荐的修复顺序
|
|
||||||
|
|
||||||
### 第一阶段:解除生产上线阻碍
|
|
||||||
|
|
||||||
1. 实现真正的限流和并发强制执行
|
|
||||||
2. 定义官方部署拓扑和工件
|
|
||||||
3. 添加 CI/CD 质量门禁
|
|
||||||
4. 加固和隔离文件处理工作负载
|
|
||||||
5. 修复后端环境配置契约
|
|
||||||
|
|
||||||
### 第二阶段:稳定运维和隐私姿态
|
|
||||||
|
|
||||||
1. 移除承载内容的日志
|
|
||||||
2. 加强请求验证
|
|
||||||
3. 深化就绪检查和指标
|
|
||||||
4. 添加前端和 E2E 自动化测试
|
|
||||||
|
|
||||||
### 第三阶段:性能和可维护性清理
|
|
||||||
|
|
||||||
1. 清理 chunk 策略
|
|
||||||
2. 限制 OCR/图像缓存
|
|
||||||
3. 移除过时代码和配置
|
|
||||||
4. 决定 PWA 支持是保留还是移除
|
|
||||||
|
|
||||||
## 最低上线门槛
|
|
||||||
|
|
||||||
至少在以下所有条件都满足之前,不应称该项目为生产就绪:
|
|
||||||
|
|
||||||
- 所有 `P0` 项目都已完成
|
|
||||||
- 日志不再捕获用户内容
|
|
||||||
- 前端和端到端自动化测试存在并在 CI 中运行
|
|
||||||
- 就绪性反映真实的上游依赖状态
|
|
||||||
- 部署和回滚已记录且可重现
|
|
||||||
- staging 环境已通过集成验证
|
|
||||||
- 至少执行了一次受控负载测试并经过审查
|
|
||||||
|
|
||||||
## 最终评估
|
|
||||||
|
|
||||||
与之前的清单相比,该项目已有实质性改进。几个严重的早期发现不再成立,特别是:
|
|
||||||
|
|
||||||
- 硬编码的认证密钥暴露
|
|
||||||
- 通配符式 CORS 姿态
|
|
||||||
- 默认公网 IP 收集
|
|
||||||
- 原始异常泄漏
|
|
||||||
|
|
||||||
然而,这一进展并不意味着已准备好投入生产。
|
|
||||||
|
|
||||||
当前仓库展示了有用的加固工作,但仍然缺乏生产服务预期的运维、测试、节流、部署和可观测性基线。
|
|
||||||
@@ -1,264 +1,93 @@
|
|||||||
# LLM in Text - 智能写作助手
|
# LLM in Text - 智能写作助手
|
||||||
|
|
||||||
基于 Vue3 和 FastAPI 的智能 Markdown 编辑器,集成大语言模型(LLM)实时补全建议功能,提供类似 GitHub Copilot 的 Ghost Text 体验。
|
基于 Vue3 和 FastAPI 的智能 Markdown 编辑器,集成大语言模型(LLM)实时补全建议功能。
|
||||||
|
|
||||||
## 功能特性
|
## 功能特性
|
||||||
|
|
||||||
### Markdown 编辑器
|
### Markdown 编辑器
|
||||||
- 基于 Milkdown Crepe 的所见即所得编辑体验
|
- 基于 Milkdown Crepe 的所见即所得编辑体验
|
||||||
- 支持完整 Markdown 语法和 LaTeX 公式
|
- 支持 Markdown 语法和 LaTeX 公式
|
||||||
|
- 支持 Mermaid 图表渲染
|
||||||
- 导入/导出 Markdown 文件
|
- 导入/导出 Markdown 文件
|
||||||
|
- 导出 DOCX 和 PDF 格式
|
||||||
|
|
||||||
### AI 智能补全
|
### AI 智能补全
|
||||||
- 实时生成文本补全建议(灰色显示)
|
- 实时生成文本补全建议(灰色显示)
|
||||||
- 流式响应,低延迟体验
|
- 流式响应,低延迟体验
|
||||||
- 多种交互方式:
|
- 多种交互方式:Tab接受、Esc拒绝、点击接受
|
||||||
- **Tab 键**:接受建议
|
|
||||||
- **Esc 键**:拒绝建议
|
|
||||||
- **点击灰色文本**:接受建议
|
|
||||||
- **继续输入**:自动拒绝建议
|
|
||||||
|
|
||||||
### AI 开关控制
|
### 文档处理
|
||||||
- 右下角 AI 开关按钮
|
- OCR 图片识别:上传图片自动识别文字
|
||||||
- 白色 = AI 启用,黑色 = AI 禁用
|
- 文档转换:PDF、DOCX、PPTX、TXT 转 Markdown
|
||||||
- 禁用时自动清除灰色文本并停止 API 调用
|
- 文档块嵌入:可折叠的文档预览块
|
||||||
|
- 智能大小限制:32KB自动禁用AI
|
||||||
|
|
||||||
|
### 设置面板
|
||||||
|
- 外观主题:亮色/暗色/跟随系统
|
||||||
|
- 背景模式:默认/暖色/阅读灯/自定义图片
|
||||||
|
- 模型智能:低/中/高思考级别
|
||||||
|
- 隐私控制:隐私模式防止发送IP
|
||||||
|
- 多语言界面:中英日韩德法
|
||||||
|
|
||||||
|
### 语音功能
|
||||||
|
- TTS文字转语音(macOS)
|
||||||
|
- STT语音转文字
|
||||||
|
|
||||||
## 技术架构
|
## 技术架构
|
||||||
|
|
||||||
```mermaid
|
前端: Vue3 + Vite + Milkdown + ProseMirror
|
||||||
flowchart TB
|
后端: FastAPI + Python + Ollama
|
||||||
subgraph Frontend["前端 (Vue3 + Vite)"]
|
|
||||||
A[App.vue] --> B[MilkdownEditor.vue]
|
|
||||||
B --> C[Crepe Editor]
|
|
||||||
C --> D[ProseMirror]
|
|
||||||
D --> E[copilotPlugin.ts]
|
|
||||||
E --> F[copilotGhostMark]
|
|
||||||
E --> G[api.js]
|
|
||||||
end
|
|
||||||
|
|
||||||
subgraph Backend["后端 (FastAPI + Python)"]
|
|
||||||
H[main.py<br/>FastAPI Server] --> I[prompt.py<br/>Prompt 构建]
|
|
||||||
H --> J[llm.py<br/>Ollama 调用]
|
|
||||||
J --> K[Ollama API]
|
|
||||||
end
|
|
||||||
|
|
||||||
G -->|POST /v1/completions<br/>SSE 流式响应| H
|
|
||||||
K -->|LLM 响应| J
|
|
||||||
```
|
|
||||||
|
|
||||||
## 项目结构
|
|
||||||
|
|
||||||
```
|
|
||||||
llm-in-text/
|
|
||||||
├── src/
|
|
||||||
│ ├── components/
|
|
||||||
│ │ └── MilkdownEditor.vue # 主编辑器组件
|
|
||||||
│ ├── plugins/
|
|
||||||
│ │ ├── copilotPlugin.ts # ProseMirror AI 补全插件
|
|
||||||
│ │ ├── types.ts # 类型定义
|
|
||||||
│ │ └── index.ts # 插件导出
|
|
||||||
│ ├── utils/
|
|
||||||
│ │ ├── api.js # API 调用封装
|
|
||||||
│ │ ├── config.js # 配置文件
|
|
||||||
│ │ └── ocrCache.js # OCR 缓存管理
|
|
||||||
│ ├── App.vue
|
|
||||||
│ └── main.js
|
|
||||||
├── backend/
|
|
||||||
│ ├── main.py # FastAPI 服务器
|
|
||||||
│ ├── llm.py # LLM API 调用
|
|
||||||
│ ├── prompt.py # Prompt 构建
|
|
||||||
│ └── requirements.txt
|
|
||||||
└── README.md
|
|
||||||
```
|
|
||||||
|
|
||||||
## 快速开始
|
## 快速开始
|
||||||
|
|
||||||
### 环境要求
|
环境: Node.js 18+、Python 3.8+、Ollama
|
||||||
- Node.js 18+
|
|
||||||
- Python 3.8+
|
|
||||||
- Ollama 服务(或其他兼容 OpenAI API 的服务)
|
|
||||||
|
|
||||||
### 安装
|
安装:
|
||||||
|
- 前端: npm install
|
||||||
|
- 后端: pip install -r backend/requirements.txt
|
||||||
|
|
||||||
```bash
|
启动:
|
||||||
# 前端
|
- 后端: python backend/main.py (端口8001)
|
||||||
npm install
|
- 前端: npm run dev (端口5173)
|
||||||
|
|
||||||
# 后端
|
## API接口
|
||||||
cd backend
|
|
||||||
pip install -r requirements.txt
|
|
||||||
```
|
|
||||||
|
|
||||||
### 配置
|
- POST /v1/completions 流式补全建议
|
||||||
|
- POST /v1/ocr 图片文字识别
|
||||||
在 `backend/.env` 中配置:
|
- POST /v1/convert 文档转换
|
||||||
|
- POST /v1/completions/cancel 取消请求
|
||||||
```env
|
|
||||||
OLLAMA_MODEL=gpt-oss:20b
|
|
||||||
OLLAMA_HOST=http://localhost:11434
|
|
||||||
```
|
|
||||||
|
|
||||||
### 启动
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# 后端(端口 8000)
|
|
||||||
cd backend
|
|
||||||
python main.py
|
|
||||||
|
|
||||||
# 前端(端口 5173)
|
|
||||||
npm run dev
|
|
||||||
```
|
|
||||||
|
|
||||||
访问 http://localhost:5173
|
|
||||||
|
|
||||||
## API 接口
|
|
||||||
|
|
||||||
### POST /v1/completions
|
|
||||||
|
|
||||||
流式获取补全建议
|
|
||||||
|
|
||||||
**请求:**
|
|
||||||
```json
|
|
||||||
{
|
|
||||||
"prefix": "# Title\n\nContent ",
|
|
||||||
"suffix": "",
|
|
||||||
"languageId": "markdown"
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
**响应(SSE):**
|
|
||||||
```
|
|
||||||
data: {"content": "here"}
|
|
||||||
data: {"content": "here is"}
|
|
||||||
data: {"done": true}
|
|
||||||
```
|
|
||||||
|
|
||||||
## 核心实现
|
## 核心实现
|
||||||
|
|
||||||
### 后端设计
|
### 后端
|
||||||
|
- main.py: FastAPI服务器、SSE流式响应
|
||||||
|
- llm.py: 异步Ollama调用、超时控制
|
||||||
|
- prompt.py: 7条Prompt规则
|
||||||
|
- tts_asr.py: macOS 语音处理
|
||||||
|
|
||||||
#### main.py - FastAPI 服务器
|
### 前端
|
||||||
- 定义 `/v1/completions` 端点
|
- copilotPlugin.ts: ProseMirror Mark系统
|
||||||
- 使用 `StreamingResponse` 返回 SSE 流式响应
|
- 关键函数: scheduleFetch、insertGhostText
|
||||||
- CORS 配置允许跨域请求
|
- Pinia Store状态管理
|
||||||
|
|
||||||
#### llm.py - LLM 调用封装
|
|
||||||
- 使用 `ollama.AsyncClient` 异步调用
|
|
||||||
- 支持 `think='high'` 思考模式
|
|
||||||
- 返回 `content` 和 `thinking` 字段
|
|
||||||
|
|
||||||
#### prompt.py - Prompt 工程
|
|
||||||
精心设计的 Prompt 模板,包含 7 条核心规则:
|
|
||||||
|
|
||||||
| 规则 | 说明 |
|
|
||||||
|------|------|
|
|
||||||
| RULE #1 | 无缝连接 - 不重复 suffix 内容,避免"复读机"错误 |
|
|
||||||
| RULE #2 | 空白处理 - 避免双空格,正确对接标点 |
|
|
||||||
| RULE #3 | 缩进对齐 - 匹配当前缩进级别和类型 |
|
|
||||||
| RULE #4 | 列表维护 - 识别并继续任务列表、有序列表、无序列表 |
|
|
||||||
| RULE #5 | 语法闭合 - 自动闭合未完成的 Markdown 语法 |
|
|
||||||
| RULE #6 | 输出格式 - 仅输出续写文本,无解释无注释 |
|
|
||||||
| RULE #7 | 必须输出 - 始终提供有用的续写建议 |
|
|
||||||
|
|
||||||
### 前端设计
|
|
||||||
|
|
||||||
#### ProseMirror Mark 系统
|
|
||||||
|
|
||||||
使用 ProseMirror 的 Mark 系统实现灰色建议文本:
|
|
||||||
|
|
||||||
```typescript
|
|
||||||
// 定义 ghost mark
|
|
||||||
export const copilotGhostMark = $markSchema('copilot_ghost', () => ({
|
|
||||||
excludes: '_',
|
|
||||||
inclusive: true,
|
|
||||||
toDOM: () => ['span', {
|
|
||||||
'data-copilot-ghost': '',
|
|
||||||
class: 'copilot-ghost-text'
|
|
||||||
}, 0]
|
|
||||||
}))
|
|
||||||
|
|
||||||
// CSS 样式
|
|
||||||
.copilot-ghost-text {
|
|
||||||
color: #999;
|
|
||||||
opacity: 0.6;
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
#### copilotPlugin 核心逻辑
|
|
||||||
|
|
||||||
```mermaid
|
|
||||||
flowchart LR
|
|
||||||
A[用户输入] --> B{文档变化?}
|
|
||||||
B -->|是| C[清除旧建议]
|
|
||||||
C --> D[防抖 1000ms]
|
|
||||||
D --> E[发送 API 请求]
|
|
||||||
E --> F[收到建议]
|
|
||||||
F --> G[插入 Ghost Text]
|
|
||||||
|
|
||||||
G --> H{用户操作}
|
|
||||||
H -->|Tab| I[接受建议<br/>移除 mark]
|
|
||||||
H -->|Esc| J[拒绝建议<br/>删除文本]
|
|
||||||
H -->|点击 Ghost| I
|
|
||||||
H -->|继续输入| J
|
|
||||||
```
|
|
||||||
|
|
||||||
#### 关键函数
|
|
||||||
|
|
||||||
| 函数 | 作用 |
|
|
||||||
|------|------|
|
|
||||||
| `scheduleFetch` | 防抖调度 API 请求 |
|
|
||||||
| `insertGhostText` | 插入带 mark 的建议文本 |
|
|
||||||
| `acceptSuggestion` | Tab 接受建议 |
|
|
||||||
| `rejectSuggestion` | Esc 拒绝建议 |
|
|
||||||
| `clearGhostText` | 清除当前建议 |
|
|
||||||
|
|
||||||
### 数据流
|
|
||||||
|
|
||||||
```mermaid
|
|
||||||
sequenceDiagram
|
|
||||||
participant U as 用户
|
|
||||||
participant E as Editor (ProseMirror)
|
|
||||||
participant P as copilotPlugin
|
|
||||||
participant A as api.js
|
|
||||||
participant B as Backend
|
|
||||||
participant L as LLM
|
|
||||||
|
|
||||||
U->>E: 输入文本
|
|
||||||
E->>P: view.update()
|
|
||||||
P->>P: 清除旧建议
|
|
||||||
P->>P: 防抖 1000ms
|
|
||||||
P->>A: fetchSuggestion(prefix, suffix)
|
|
||||||
A->>B: POST /v1/completions
|
|
||||||
B->>B: build_prompt()
|
|
||||||
B->>L: ollama.chat()
|
|
||||||
L-->>B: {content, thinking}
|
|
||||||
B-->>A: SSE stream
|
|
||||||
A-->>P: suggestion text
|
|
||||||
P->>E: insertGhostText()
|
|
||||||
E-->>U: 显示灰色建议
|
|
||||||
|
|
||||||
alt Tab 键
|
|
||||||
U->>P: Tab
|
|
||||||
P->>E: acceptSuggestion()
|
|
||||||
E-->>U: 建议变为正常文本
|
|
||||||
else Esc 键
|
|
||||||
U->>P: Esc
|
|
||||||
P->>E: rejectSuggestion()
|
|
||||||
E-->>U: 建议消失
|
|
||||||
else 继续输入
|
|
||||||
U->>E: 输入其他字符
|
|
||||||
E->>P: handleKeyDown()
|
|
||||||
P->>E: clearGhostText()
|
|
||||||
end
|
|
||||||
```
|
|
||||||
|
|
||||||
## 设计亮点
|
## 设计亮点
|
||||||
|
|
||||||
1. **前后端分离**:前端只负责渲染和数据回传,后端负责 LLM 调用、Prompt 构建和数据解析
|
1. 前后端分离
|
||||||
2. **低延迟优化**:防抖机制 (1000ms) + SSE 流式响应 + AbortController 取消过期请求
|
2. 低延迟优化:防抖+SSE+AbortController
|
||||||
3. **ProseMirror Mark 系统**:与编辑器状态完美集成,支持 Undo/Redo
|
3. ProseMirror Mark系统
|
||||||
4. **多种交互方式**:Tab/Esc/点击/输入,用户体验友好
|
4. 多种交互方式
|
||||||
5. **智能大小限制**:文档超过 32KB 自动禁用 AI 功能
|
5. 智能大小限制
|
||||||
|
6. 隐私保护
|
||||||
|
7. 多语言支持
|
||||||
|
8. 主题定制
|
||||||
|
9. 文档处理
|
||||||
|
10. 语音功能
|
||||||
|
|
||||||
|
## 开发指南
|
||||||
|
|
||||||
|
代码风格: Python(4空格,snake_case) JS/TS(2空格,camelCase)
|
||||||
|
测试: pytest
|
||||||
|
构建: npm run build
|
||||||
|
|
||||||
## 许可证
|
## 许可证
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,20 @@
|
|||||||
|
const path = require('path')
|
||||||
|
const { convert } = require('docx2pdf-converter')
|
||||||
|
|
||||||
|
function main() {
|
||||||
|
const inputPath = process.argv[2]
|
||||||
|
const outputPath = process.argv[3]
|
||||||
|
|
||||||
|
if (!inputPath || !outputPath) {
|
||||||
|
throw new Error('缺少 DOCX 或 PDF 路径')
|
||||||
|
}
|
||||||
|
|
||||||
|
convert(path.resolve(inputPath), path.resolve(outputPath))
|
||||||
|
}
|
||||||
|
|
||||||
|
try {
|
||||||
|
main()
|
||||||
|
} catch (error) {
|
||||||
|
console.error(error instanceof Error ? error.message : String(error))
|
||||||
|
process.exit(1)
|
||||||
|
}
|
||||||
+2
-2
@@ -63,10 +63,10 @@ def _extract_message(response) -> tuple[str, str]:
|
|||||||
async def call_ollama(
|
async def call_ollama(
|
||||||
prompt: str,
|
prompt: str,
|
||||||
*,
|
*,
|
||||||
system_prompt: str = None,
|
system_prompt: str | None = None,
|
||||||
tag: str = "default",
|
tag: str = "default",
|
||||||
temperature: float = 0.7,
|
temperature: float = 0.7,
|
||||||
thinking: str = None,
|
thinking: str | None = None,
|
||||||
) -> dict:
|
) -> dict:
|
||||||
"""
|
"""
|
||||||
调用 Ollama API 并返回 content 和 thinking。
|
调用 Ollama API 并返回 content 和 thinking。
|
||||||
|
|||||||
+156
-113
@@ -1,17 +1,21 @@
|
|||||||
import asyncio
|
import asyncio
|
||||||
import base64
|
import base64
|
||||||
import json
|
|
||||||
import logging
|
import logging
|
||||||
import os
|
import os
|
||||||
|
import re
|
||||||
|
import shutil
|
||||||
|
import subprocess
|
||||||
import tempfile
|
import tempfile
|
||||||
import uuid
|
import uuid
|
||||||
from typing import Optional
|
from typing import Optional
|
||||||
|
|
||||||
from fastapi import FastAPI, Request
|
from fastapi import FastAPI, HTTPException, Request, Security, File, UploadFile
|
||||||
from fastapi.middleware.cors import CORSMiddleware
|
from fastapi.middleware.cors import CORSMiddleware
|
||||||
from fastapi.responses import JSONResponse, StreamingResponse
|
from fastapi.responses import JSONResponse, Response
|
||||||
|
from fastapi.security import APIKeyHeader
|
||||||
from pydantic import BaseModel
|
from pydantic import BaseModel
|
||||||
|
|
||||||
|
from geoip import get_ip_location_text
|
||||||
from llm import call_ollama, call_vlm_ocr
|
from llm import call_ollama, call_vlm_ocr
|
||||||
from prompt import build_completion_prompts, prepare_prompt_context
|
from prompt import build_completion_prompts, prepare_prompt_context
|
||||||
import markitdown
|
import markitdown
|
||||||
@@ -24,33 +28,28 @@ logger = logging.getLogger("api")
|
|||||||
|
|
||||||
app = FastAPI()
|
app = FastAPI()
|
||||||
|
|
||||||
app.add_middleware(
|
|
||||||
CORSMiddleware,
|
|
||||||
allow_origins=[
|
|
||||||
"http://localhost:5173",
|
|
||||||
"http://localhost:3000",
|
|
||||||
"https://www.imageteach.tech",
|
|
||||||
"https://chat.imageteach.tech",
|
|
||||||
],
|
|
||||||
allow_credentials=False,
|
|
||||||
allow_methods=["POST", "OPTIONS"],
|
|
||||||
allow_headers=["Content-Type", "X-Request-Id"],
|
|
||||||
)
|
|
||||||
|
|
||||||
ACTIVE_COMPLETIONS: dict[str, asyncio.Task] = {}
|
ACTIVE_COMPLETIONS: dict[str, asyncio.Task] = {}
|
||||||
ACTIVE_COMPLETIONS_LOCK = asyncio.Lock()
|
ACTIVE_COMPLETIONS_LOCK = asyncio.Lock()
|
||||||
|
|
||||||
# Rate limiting
|
app.add_middleware(
|
||||||
MAX_CONCURRENT_COMPLETIONS = 4
|
CORSMiddleware,
|
||||||
COMPLETION_RATE_LIMIT = 60 # per minute
|
allow_origins=["*"],
|
||||||
|
allow_credentials=True,
|
||||||
|
allow_methods=["*"],
|
||||||
|
allow_headers=["*", "X-API-Key", "X-Client-IP", "X-Request-Id"],
|
||||||
|
)
|
||||||
|
|
||||||
# File size limits (bytes)
|
API_KEY = "your-secret-key-here"
|
||||||
MAX_IMAGE_SIZE = 10 * 1024 * 1024 # 10MB
|
api_key_header = APIKeyHeader(name="X-API-Key")
|
||||||
MAX_CONVERT_SIZE = 50 * 1024 * 1024 # 50MB
|
|
||||||
|
|
||||||
# Allowed file extensions
|
|
||||||
ALLOWED_IMAGE_EXTENSIONS = {".jpg", ".jpeg", ".png", ".webp"}
|
async def get_api_key(api_key: str = Security(api_key_header)):
|
||||||
ALLOWED_CONVERT_EXTENSIONS = {".pdf", ".docx", ".pptx", ".xlsx", ".md", ".txt"}
|
if api_key != API_KEY:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=403,
|
||||||
|
detail="Could not validate credentials",
|
||||||
|
)
|
||||||
|
return api_key
|
||||||
|
|
||||||
|
|
||||||
class UserPreferences(BaseModel):
|
class UserPreferences(BaseModel):
|
||||||
@@ -84,6 +83,32 @@ class ConvertRequest(BaseModel):
|
|||||||
filename: str = "document.pdf"
|
filename: str = "document.pdf"
|
||||||
|
|
||||||
|
|
||||||
|
ALLOWED_CONVERT_EXTENSIONS = {".txt", ".docx", ".pptx", ".pdf"}
|
||||||
|
IMAGE_MARKDOWN_RE = re.compile(r"!\[[^\]]*]\([^)]+\)")
|
||||||
|
IMAGE_HTML_RE = re.compile(r"<img\b[^>]*>", re.IGNORECASE)
|
||||||
|
|
||||||
|
|
||||||
|
def _convert_docx_to_pdf(input_path: str, output_path: str) -> None:
|
||||||
|
node_executable = shutil.which("node")
|
||||||
|
if not node_executable:
|
||||||
|
raise RuntimeError("未找到 Node.js,无法转换 DOCX 为 PDF")
|
||||||
|
|
||||||
|
bridge_path = os.path.join(os.path.dirname(__file__), "docx2pdf_bridge.cjs")
|
||||||
|
if not os.path.exists(bridge_path):
|
||||||
|
raise RuntimeError("缺少 DOCX 转 PDF 桥接脚本")
|
||||||
|
|
||||||
|
result = subprocess.run(
|
||||||
|
[node_executable, bridge_path, input_path, output_path],
|
||||||
|
cwd=os.path.dirname(os.path.dirname(__file__)),
|
||||||
|
capture_output=True,
|
||||||
|
text=True,
|
||||||
|
)
|
||||||
|
|
||||||
|
if result.returncode != 0:
|
||||||
|
error_text = (result.stderr or result.stdout or "DOCX 转 PDF 失败").strip()
|
||||||
|
raise RuntimeError(error_text)
|
||||||
|
|
||||||
|
|
||||||
def _preview(text: str, limit: int = 80) -> str:
|
def _preview(text: str, limit: int = 80) -> str:
|
||||||
value = (text or "").replace("\n", "\\n")
|
value = (text or "").replace("\n", "\\n")
|
||||||
if len(value) <= limit:
|
if len(value) <= limit:
|
||||||
@@ -91,34 +116,41 @@ def _preview(text: str, limit: int = 80) -> str:
|
|||||||
return value[:limit] + "..."
|
return value[:limit] + "..."
|
||||||
|
|
||||||
|
|
||||||
def _error_response(request_id: str, code: str, message: str, status_code: int = 500) -> JSONResponse:
|
def _sanitize_converted_markdown(text: str) -> str:
|
||||||
return JSONResponse(
|
value = (text or "").replace("\r\n", "\n").replace("\r", "\n")
|
||||||
content={
|
value = IMAGE_MARKDOWN_RE.sub("", value)
|
||||||
"error": {
|
value = IMAGE_HTML_RE.sub("", value)
|
||||||
"code": code,
|
value = re.sub(r"\n{3,}", "\n\n", value)
|
||||||
"message": message,
|
return value.strip()
|
||||||
"request_id": request_id,
|
|
||||||
}
|
|
||||||
},
|
|
||||||
status_code=status_code,
|
|
||||||
)
|
|
||||||
|
|
||||||
|
|
||||||
def _sse_payload(payload: dict) -> str:
|
def get_client_ip(request: Request) -> str:
|
||||||
return f"data: {json.dumps(payload)}\n\n"
|
if request.client:
|
||||||
|
return request.headers.get("X-Client-IP") or request.client.host
|
||||||
|
return request.headers.get("X-Client-IP") or "unknown"
|
||||||
|
|
||||||
|
|
||||||
@app.post("/v1/completions")
|
@app.post("/v1/completions")
|
||||||
async def create_completion(request: Request, req: CompletionRequest):
|
async def create_completion(request: Request, req: CompletionRequest, api_key: str = Security(get_api_key)):
|
||||||
request_id = request.headers.get("X-Request-Id") or str(uuid.uuid4())
|
request_id = request.headers.get("X-Request-Id") or str(uuid.uuid4())
|
||||||
request_tag = request_id[:8]
|
request_tag = request_id[:8]
|
||||||
inference_task: Optional[asyncio.Task] = None
|
inference_task: Optional[asyncio.Task] = None
|
||||||
|
|
||||||
|
client_ip = "hidden"
|
||||||
|
location = ""
|
||||||
|
|
||||||
|
if not req.privacy_mode:
|
||||||
|
client_ip = get_client_ip(request)
|
||||||
|
location = get_ip_location_text(client_ip)
|
||||||
|
if location:
|
||||||
|
logger.info("[%s] client_location=%s", request_tag, location)
|
||||||
|
|
||||||
try:
|
try:
|
||||||
logger.info(
|
logger.info(
|
||||||
"[%s] /v1/completions request_id=%s prefix_chars=%d suffix_chars=%d lang=%s thinking=%s privacy=%s",
|
"[%s] /v1/completions request_id=%s client_ip=%s prefix_chars=%d suffix_chars=%d lang=%s thinking=%s privacy=%s",
|
||||||
request_tag,
|
request_tag,
|
||||||
request_id,
|
request_id,
|
||||||
|
client_ip,
|
||||||
len(req.prefix or ""),
|
len(req.prefix or ""),
|
||||||
len(req.suffix or ""),
|
len(req.suffix or ""),
|
||||||
req.languageId,
|
req.languageId,
|
||||||
@@ -134,6 +166,7 @@ async def create_completion(request: Request, req: CompletionRequest):
|
|||||||
req.prefix,
|
req.prefix,
|
||||||
req.suffix,
|
req.suffix,
|
||||||
req.languageId,
|
req.languageId,
|
||||||
|
location=location,
|
||||||
thinking_level=req.model_thinking,
|
thinking_level=req.model_thinking,
|
||||||
preferences=req.user_preferences,
|
preferences=req.user_preferences,
|
||||||
)
|
)
|
||||||
@@ -148,11 +181,10 @@ async def create_completion(request: Request, req: CompletionRequest):
|
|||||||
)
|
)
|
||||||
)
|
)
|
||||||
|
|
||||||
async with ACTIVE_COMPLETIONS_LOCK:
|
existing = ACTIVE_COMPLETIONS.get(request_id)
|
||||||
existing = ACTIVE_COMPLETIONS.get(request_id)
|
if existing and not existing.done():
|
||||||
if existing and not existing.done():
|
existing.cancel()
|
||||||
existing.cancel()
|
ACTIVE_COMPLETIONS[request_id] = inference_task
|
||||||
ACTIVE_COMPLETIONS[request_id] = inference_task
|
|
||||||
|
|
||||||
result = await inference_task
|
result = await inference_task
|
||||||
content = result["content"] or ""
|
content = result["content"] or ""
|
||||||
@@ -166,30 +198,21 @@ async def create_completion(request: Request, req: CompletionRequest):
|
|||||||
_preview(content, 120),
|
_preview(content, 120),
|
||||||
)
|
)
|
||||||
|
|
||||||
async def generate():
|
return JSONResponse(content={"content": content, "request_id": request_id})
|
||||||
yield _sse_payload({"content": content})
|
|
||||||
yield _sse_payload({"done": True})
|
|
||||||
|
|
||||||
return StreamingResponse(generate(), media_type="text/event-stream")
|
|
||||||
except asyncio.CancelledError:
|
except asyncio.CancelledError:
|
||||||
logger.info("[%s] /v1/completions cancelled request_id=%s", request_tag, request_id)
|
logger.info("[%s] /v1/completions cancelled request_id=%s", request_tag, request_id)
|
||||||
|
return JSONResponse(content={"cancelled": True, "request_id": request_id}, status_code=499)
|
||||||
async def cancelled():
|
|
||||||
yield _sse_payload({"cancelled": True, "request_id": request_id, "done": True})
|
|
||||||
|
|
||||||
return StreamingResponse(cancelled(), media_type="text/event-stream")
|
|
||||||
except Exception as e:
|
except Exception as e:
|
||||||
logger.exception("[%s] /v1/completions failed request_id=%s: %s", request_tag, request_id, e)
|
logger.exception("[%s] /v1/completions failed request_id=%s: %s", request_tag, request_id, e)
|
||||||
return _error_response(request_id, "INTERNAL_ERROR", "Service temporarily unavailable", 500)
|
return JSONResponse(content={"error": str(e)}, status_code=500)
|
||||||
finally:
|
finally:
|
||||||
async with ACTIVE_COMPLETIONS_LOCK:
|
active = ACTIVE_COMPLETIONS.get(request_id)
|
||||||
active = ACTIVE_COMPLETIONS.get(request_id)
|
if active is not None and active is inference_task:
|
||||||
if active is not None and active is inference_task:
|
ACTIVE_COMPLETIONS.pop(request_id, None)
|
||||||
ACTIVE_COMPLETIONS.pop(request_id, None)
|
|
||||||
|
|
||||||
|
|
||||||
@app.post("/v1/completions/cancel")
|
@app.post("/v1/completions/cancel")
|
||||||
async def cancel_completion(req: CancelCompletionRequest):
|
async def cancel_completion(req: CancelCompletionRequest, api_key: str = Security(get_api_key)):
|
||||||
request_tag = str(uuid.uuid4())[:8]
|
request_tag = str(uuid.uuid4())[:8]
|
||||||
request_id = req.request_id or ""
|
request_id = req.request_id or ""
|
||||||
|
|
||||||
@@ -225,7 +248,7 @@ async def cancel_completion(req: CancelCompletionRequest):
|
|||||||
|
|
||||||
|
|
||||||
@app.post("/v1/ocr")
|
@app.post("/v1/ocr")
|
||||||
async def ocr_image(request: OCRRequest):
|
async def ocr_image(request: OCRRequest, api_key: str = Security(get_api_key)):
|
||||||
request_id = str(uuid.uuid4())[:8]
|
request_id = str(uuid.uuid4())[:8]
|
||||||
try:
|
try:
|
||||||
logger.info(
|
logger.info(
|
||||||
@@ -235,22 +258,7 @@ async def ocr_image(request: OCRRequest):
|
|||||||
request.language,
|
request.language,
|
||||||
len(request.image or ""),
|
len(request.image or ""),
|
||||||
)
|
)
|
||||||
|
|
||||||
# Check file size before decoding
|
|
||||||
if len(request.image or "") > MAX_IMAGE_SIZE * 4 // 3: # base64 overhead
|
|
||||||
return _error_response(request_id, "FILE_TOO_LARGE", "Image exceeds 10MB limit", 413)
|
|
||||||
|
|
||||||
# Check extension
|
|
||||||
ext = os.path.splitext(request.filename)[1].lower()
|
|
||||||
if ext not in ALLOWED_IMAGE_EXTENSIONS:
|
|
||||||
return _error_response(request_id, "INVALID_FILE_TYPE", "Only jpg/png/webp allowed", 415)
|
|
||||||
|
|
||||||
image_bytes = base64.b64decode(request.image)
|
image_bytes = base64.b64decode(request.image)
|
||||||
|
|
||||||
# Check actual decoded size
|
|
||||||
if len(image_bytes) > MAX_IMAGE_SIZE:
|
|
||||||
return _error_response(request_id, "FILE_TOO_LARGE", "Image exceeds 10MB limit", 413)
|
|
||||||
|
|
||||||
logger.info("[%s] /v1/ocr decoded image_bytes=%d", request_id, len(image_bytes))
|
logger.info("[%s] /v1/ocr decoded image_bytes=%d", request_id, len(image_bytes))
|
||||||
result = await call_vlm_ocr(image_bytes, request.language)
|
result = await call_vlm_ocr(image_bytes, request.language)
|
||||||
logger.info(
|
logger.info(
|
||||||
@@ -262,12 +270,12 @@ async def ocr_image(request: OCRRequest):
|
|||||||
return {"text": result, "filename": request.filename}
|
return {"text": result, "filename": request.filename}
|
||||||
except Exception as e:
|
except Exception as e:
|
||||||
logger.exception("[%s] /v1/ocr failed: %s", request_id, e)
|
logger.exception("[%s] /v1/ocr failed: %s", request_id, e)
|
||||||
return _error_response(request_id, "OCR_FAILED", "Failed to process image", 500)
|
return JSONResponse(content={"error": str(e)}, status_code=500)
|
||||||
|
|
||||||
|
|
||||||
@app.post("/v1/convert")
|
@app.post("/v1/convert")
|
||||||
async def convert_to_markdown(request: ConvertRequest):
|
async def convert_to_markdown(request: ConvertRequest, api_key: str = Security(get_api_key)):
|
||||||
"""将文件转换为Markdown格式"""
|
"""Convert file to markdown"""
|
||||||
request_id = str(uuid.uuid4())[:8]
|
request_id = str(uuid.uuid4())[:8]
|
||||||
|
|
||||||
try:
|
try:
|
||||||
@@ -278,34 +286,33 @@ async def convert_to_markdown(request: ConvertRequest):
|
|||||||
len(request.file or ""),
|
len(request.file or ""),
|
||||||
)
|
)
|
||||||
|
|
||||||
# Check file size before decoding
|
# Decode base64
|
||||||
if len(request.file or "") > MAX_CONVERT_SIZE * 4 // 3:
|
|
||||||
return _error_response(request_id, "FILE_TOO_LARGE", "File exceeds 50MB limit", 413)
|
|
||||||
|
|
||||||
# Get file extension and validate
|
|
||||||
ext = os.path.splitext(request.filename)[1].lower()
|
|
||||||
if ext not in ALLOWED_CONVERT_EXTENSIONS:
|
|
||||||
return _error_response(request_id, "INVALID_FILE_TYPE", "Only pdf/docx/pptx/xlsx/md/txt allowed", 415)
|
|
||||||
|
|
||||||
# 解码Base64文件内容
|
|
||||||
file_bytes = base64.b64decode(request.file)
|
file_bytes = base64.b64decode(request.file)
|
||||||
|
|
||||||
# Check actual decoded size
|
|
||||||
if len(file_bytes) > MAX_CONVERT_SIZE:
|
|
||||||
return _error_response(request_id, "FILE_TOO_LARGE", "File exceeds 50MB limit", 413)
|
|
||||||
|
|
||||||
logger.info("[%s] /v1/convert decoded file_bytes=%d", request_id, len(file_bytes))
|
logger.info("[%s] /v1/convert decoded file_bytes=%d", request_id, len(file_bytes))
|
||||||
|
|
||||||
# 创建临时文件
|
# Get file extension
|
||||||
|
ext = os.path.splitext(request.filename)[1].lower()
|
||||||
|
|
||||||
|
if ext not in ALLOWED_CONVERT_EXTENSIONS:
|
||||||
|
raise ValueError("仅支持 txt、docx、pptx、pdf 格式")
|
||||||
|
|
||||||
|
if ext == ".txt":
|
||||||
|
markdown_text = _sanitize_converted_markdown(file_bytes.decode("utf-8", errors="ignore"))
|
||||||
|
return {
|
||||||
|
"markdown": markdown_text,
|
||||||
|
"filename": request.filename
|
||||||
|
}
|
||||||
|
|
||||||
|
# Create temporary file
|
||||||
with tempfile.NamedTemporaryFile(delete=False, suffix=ext) as tmp:
|
with tempfile.NamedTemporaryFile(delete=False, suffix=ext) as tmp:
|
||||||
tmp.write(file_bytes)
|
tmp.write(file_bytes)
|
||||||
tmp_path = tmp.name
|
tmp_path = tmp.name
|
||||||
|
|
||||||
try:
|
try:
|
||||||
# 使用MarkItDown转换为Markdown
|
# Convert using MarkItDown
|
||||||
md = markitdown.MarkItDown()
|
md = markitdown.MarkItDown()
|
||||||
result = md.convert(tmp_path)
|
result = md.convert(tmp_path)
|
||||||
markdown_text = result.text_content
|
markdown_text = _sanitize_converted_markdown(result.text_content)
|
||||||
|
|
||||||
logger.info(
|
logger.info(
|
||||||
"[%s] /v1/convert success text_chars=%d text_preview='%s'",
|
"[%s] /v1/convert success text_chars=%d text_preview='%s'",
|
||||||
@@ -319,13 +326,56 @@ async def convert_to_markdown(request: ConvertRequest):
|
|||||||
"filename": request.filename
|
"filename": request.filename
|
||||||
}
|
}
|
||||||
finally:
|
finally:
|
||||||
# 清理临时文件
|
# Clean up temporary file
|
||||||
if os.path.exists(tmp_path):
|
if os.path.exists(tmp_path):
|
||||||
os.unlink(tmp_path)
|
os.unlink(tmp_path)
|
||||||
|
|
||||||
except Exception as e:
|
except Exception as e:
|
||||||
logger.exception("[%s] /v1/convert failed: %s", request_id, e)
|
logger.exception("[%s] /v1/convert failed: %s", request_id, e)
|
||||||
return _error_response(request_id, "CONVERT_FAILED", "Failed to convert file", 500)
|
return JSONResponse(content={"error": str(e)}, status_code=500)
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
@app.post("/v1/export/pdf")
|
||||||
|
async def export_pdf(file: UploadFile = File(...), api_key: str = Security(get_api_key)):
|
||||||
|
request_id = str(uuid.uuid4())[:8]
|
||||||
|
original_name = file.filename or "document.docx"
|
||||||
|
base_name = os.path.splitext(original_name)[0] or "document"
|
||||||
|
|
||||||
|
try:
|
||||||
|
file_bytes = await file.read()
|
||||||
|
logger.info(
|
||||||
|
"[%s] /v1/export/pdf filename=%s file_bytes=%d",
|
||||||
|
request_id,
|
||||||
|
original_name,
|
||||||
|
len(file_bytes),
|
||||||
|
)
|
||||||
|
|
||||||
|
with tempfile.TemporaryDirectory() as temp_dir:
|
||||||
|
input_path = os.path.join(temp_dir, f"{base_name}.docx")
|
||||||
|
output_path = os.path.join(temp_dir, f"{base_name}.pdf")
|
||||||
|
|
||||||
|
with open(input_path, "wb") as tmp_file:
|
||||||
|
tmp_file.write(file_bytes)
|
||||||
|
|
||||||
|
await asyncio.to_thread(_convert_docx_to_pdf, input_path, output_path)
|
||||||
|
|
||||||
|
if not os.path.exists(output_path):
|
||||||
|
raise RuntimeError("PDF 转换后未生成输出文件")
|
||||||
|
|
||||||
|
with open(output_path, "rb") as pdf_file:
|
||||||
|
pdf_bytes = pdf_file.read()
|
||||||
|
|
||||||
|
logger.info("[%s] /v1/export/pdf success pdf_bytes=%d", request_id, len(pdf_bytes))
|
||||||
|
headers = {
|
||||||
|
"Content-Disposition": f'attachment; filename="{base_name}.pdf"',
|
||||||
|
}
|
||||||
|
return Response(content=pdf_bytes, media_type="application/pdf", headers=headers)
|
||||||
|
except Exception as e:
|
||||||
|
logger.exception("[%s] /v1/export/pdf failed: %s", request_id, e)
|
||||||
|
return JSONResponse(content={"error": str(e)}, status_code=500)
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
@@ -334,17 +384,10 @@ if __name__ == "__main__":
|
|||||||
uvicorn.run(app, host="0.0.0.0", port=8001)
|
uvicorn.run(app, host="0.0.0.0", port=8001)
|
||||||
|
|
||||||
|
|
||||||
@app.get("/health/live")
|
# TTS and ASR routes (lazy loaded to avoid heavy import on startup)
|
||||||
async def health_live():
|
def _register_tts_asr_routes():
|
||||||
return {"status": "ok"}
|
from tts_asr import register_tts_asr_routes
|
||||||
|
register_tts_asr_routes(app)
|
||||||
|
|
||||||
|
_register_tts_asr_routes()
|
||||||
|
|
||||||
@app.get("/health/ready")
|
|
||||||
async def health_ready():
|
|
||||||
# Check if critical components are available
|
|
||||||
try:
|
|
||||||
# Could add more checks here (e.g., Ollama connectivity)
|
|
||||||
return {"status": "ready"}
|
|
||||||
except Exception as e:
|
|
||||||
logger.warning("[health/ready] not ready: %s", e)
|
|
||||||
return _error_response("health-check", "NOT_READY", "Service not ready", 503)
|
|
||||||
|
|||||||
+258
-199
@@ -1,6 +1,13 @@
|
|||||||
from datetime import datetime, timedelta, timezone
|
from datetime import datetime, timedelta, timezone
|
||||||
import re
|
import re
|
||||||
from typing import Tuple
|
from typing import Protocol, Tuple, runtime_checkable
|
||||||
|
|
||||||
|
|
||||||
|
@runtime_checkable
|
||||||
|
class UserPreferences(Protocol):
|
||||||
|
language: str
|
||||||
|
currency: str
|
||||||
|
timezone: str
|
||||||
|
|
||||||
|
|
||||||
def _get_current_datetime(timezone_pref: str = "auto") -> str:
|
def _get_current_datetime(timezone_pref: str = "auto") -> str:
|
||||||
@@ -214,113 +221,107 @@ def _canonical_language_id(language_id: str) -> str:
|
|||||||
return LANGUAGE_SYNONYMS.get(safe, safe)
|
return LANGUAGE_SYNONYMS.get(safe, safe)
|
||||||
|
|
||||||
|
|
||||||
def _language_guidance(language_id: str) -> str:
|
_JS_LANGS = {"javascript", "typescript"}
|
||||||
canonical = _canonical_language_id(language_id)
|
_CODE_LANGS = {"python", "go", "rust", "java", "kotlin", "swift", "ruby", "php", "lua", "c", "cpp", "csharp", "r", "matlab", "dart"}
|
||||||
if canonical == "markdown":
|
|
||||||
return ""
|
_LANG_GUIDANCE = {
|
||||||
if canonical == "mermaid":
|
"mermaid": """
|
||||||
return """
|
|
||||||
Language-specific guidance (mermaid):
|
Language-specific guidance (mermaid):
|
||||||
- Output valid Mermaid syntax only.
|
- Output valid Mermaid syntax only.
|
||||||
- Prefer concise, syntactically correct diagram statements.
|
- Prefer concise, syntactically correct diagram statements.
|
||||||
- Avoid prose unless the user prompt explicitly requires it."""
|
- Avoid prose unless the user prompt explicitly requires it.""",
|
||||||
if canonical == "latex":
|
"latex": """
|
||||||
return """
|
|
||||||
Language-specific guidance (latex):
|
Language-specific guidance (latex):
|
||||||
- Output LaTeX math content only when completing LaTeX.
|
- Output LaTeX math content only when completing LaTeX.
|
||||||
- If CURSOR_IN_FENCED_CODE_BLOCK=true and CURSOR_FENCE_LANGUAGE is latex/tex/katex:
|
- If CURSOR_IN_FENCED_CODE_BLOCK=true and CURSOR_FENCE_LANGUAGE is latex/tex/katex:
|
||||||
- Output raw LaTeX lines only.
|
- Output raw LaTeX lines only.
|
||||||
- Do not wrap with $ or $$."""
|
- Do not wrap with $ or $$.""",
|
||||||
if canonical == "json":
|
"json": """
|
||||||
return """
|
|
||||||
Language-specific guidance (json):
|
Language-specific guidance (json):
|
||||||
- Output strict JSON only (no comments, no trailing commas).
|
- Output strict JSON only (no comments, no trailing commas).
|
||||||
- Ensure valid quotes and braces."""
|
- Ensure valid quotes and braces.""",
|
||||||
if canonical == "yaml":
|
"yaml": """
|
||||||
return """
|
|
||||||
Language-specific guidance (yaml):
|
Language-specific guidance (yaml):
|
||||||
- Output valid YAML only.
|
- Output valid YAML only.
|
||||||
- Use consistent indentation and avoid tabs."""
|
- Use consistent indentation and avoid tabs.""",
|
||||||
if canonical == "toml":
|
"toml": """
|
||||||
return """
|
|
||||||
Language-specific guidance (toml):
|
Language-specific guidance (toml):
|
||||||
- Output valid TOML only.
|
- Output valid TOML only.
|
||||||
- Keep key types consistent."""
|
- Keep key types consistent.""",
|
||||||
if canonical == "ini":
|
"ini": """
|
||||||
return """
|
|
||||||
Language-specific guidance (ini):
|
Language-specific guidance (ini):
|
||||||
- Output valid INI only.
|
- Output valid INI only.
|
||||||
- Keep section headers and key=value pairs consistent."""
|
- Keep section headers and key=value pairs consistent.""",
|
||||||
if canonical == "sql":
|
"sql": """
|
||||||
return """
|
|
||||||
Language-specific guidance (sql):
|
Language-specific guidance (sql):
|
||||||
- Output a single, valid SQL statement unless context requires multiple.
|
- Output a single, valid SQL statement unless context requires multiple.
|
||||||
- Prefer ANSI SQL when dialect is unclear."""
|
- Prefer ANSI SQL when dialect is unclear.""",
|
||||||
if canonical == "bash":
|
"bash": """
|
||||||
return """
|
|
||||||
Language-specific guidance (bash):
|
Language-specific guidance (bash):
|
||||||
- Output POSIX-compatible shell when possible.
|
- Output POSIX-compatible shell when possible.
|
||||||
- Avoid interactive prompts or destructive commands unless requested."""
|
- Avoid interactive prompts or destructive commands unless requested.""",
|
||||||
if canonical == "powershell":
|
"powershell": """
|
||||||
return """
|
|
||||||
Language-specific guidance (powershell):
|
Language-specific guidance (powershell):
|
||||||
- Output valid PowerShell commands.
|
- Output valid PowerShell commands.
|
||||||
- Avoid destructive commands unless explicitly requested."""
|
- Avoid destructive commands unless explicitly requested.""",
|
||||||
if canonical == "html":
|
"html": """
|
||||||
return """
|
|
||||||
Language-specific guidance (html):
|
Language-specific guidance (html):
|
||||||
- Output valid HTML only.
|
- Output valid HTML only.
|
||||||
- Keep markup minimal and well-formed."""
|
- Keep markup minimal and well-formed.""",
|
||||||
if canonical == "css":
|
"css": """
|
||||||
return """
|
|
||||||
Language-specific guidance (css):
|
Language-specific guidance (css):
|
||||||
- Output valid CSS only.
|
- Output valid CSS only.
|
||||||
- Use concise, readable selectors."""
|
- Use concise, readable selectors.""",
|
||||||
if canonical == "diff":
|
"diff": """
|
||||||
return """
|
|
||||||
Language-specific guidance (diff):
|
Language-specific guidance (diff):
|
||||||
- Output a unified diff only.
|
- Output a unified diff only.
|
||||||
- Ensure @@ hunk headers and +/- lines are consistent."""
|
- Ensure @@ hunk headers and +/- lines are consistent.""",
|
||||||
if canonical == "regex":
|
"regex": """
|
||||||
return """
|
|
||||||
Language-specific guidance (regex):
|
Language-specific guidance (regex):
|
||||||
- Output the regex pattern only.
|
- Output the regex pattern only.
|
||||||
- Avoid delimiters unless explicitly requested."""
|
- Avoid delimiters unless explicitly requested.""",
|
||||||
if canonical in {"javascript", "typescript"}:
|
"text": """
|
||||||
return f"""
|
|
||||||
Language-specific guidance ({canonical}):
|
|
||||||
- Output valid {canonical} code.
|
|
||||||
- Prefer modern syntax and avoid prose unless comments are needed."""
|
|
||||||
if canonical in {"python", "go", "rust", "java", "kotlin", "swift", "ruby", "php", "lua", "c", "cpp", "csharp", "r", "matlab", "dart"}:
|
|
||||||
return f"""
|
|
||||||
Language-specific guidance ({canonical}):
|
|
||||||
- Output valid {canonical} code.
|
|
||||||
- Avoid prose unless context clearly expects comments or docstrings."""
|
|
||||||
if canonical == "text":
|
|
||||||
return """
|
|
||||||
Language-specific guidance (text):
|
Language-specific guidance (text):
|
||||||
- Output plain text only.
|
- Output plain text only.
|
||||||
- Avoid markdown formatting unless explicitly asked."""
|
- Avoid markdown formatting unless explicitly asked.""",
|
||||||
if canonical == "xml":
|
"xml": """
|
||||||
return """
|
|
||||||
Language-specific guidance (xml):
|
Language-specific guidance (xml):
|
||||||
- Output well-formed XML only.
|
- Output well-formed XML only.
|
||||||
- Ensure matching tags and proper escaping."""
|
- Ensure matching tags and proper escaping.""",
|
||||||
if canonical == "dockerfile":
|
"dockerfile": """
|
||||||
return """
|
|
||||||
Language-specific guidance (dockerfile):
|
Language-specific guidance (dockerfile):
|
||||||
- Output valid Dockerfile instructions only.
|
- Output valid Dockerfile instructions only.
|
||||||
- Keep layers minimal and ordered logically."""
|
- Keep layers minimal and ordered logically.""",
|
||||||
if canonical == "makefile":
|
"makefile": """
|
||||||
return """
|
|
||||||
Language-specific guidance (makefile):
|
Language-specific guidance (makefile):
|
||||||
- Output valid Makefile syntax only.
|
- Output valid Makefile syntax only.
|
||||||
- Use tabs for recipe lines."""
|
- Use tabs for recipe lines.""",
|
||||||
return f"""
|
}
|
||||||
Language-specific guidance ({canonical}):
|
|
||||||
- Output valid {canonical} code.
|
_GENERIC_CODE = """
|
||||||
|
Language-specific guidance ({lang}):
|
||||||
|
- Output valid {lang} code.
|
||||||
- Avoid prose unless context clearly expects comments or docstrings."""
|
- Avoid prose unless context clearly expects comments or docstrings."""
|
||||||
|
|
||||||
|
_JS_CODE = """
|
||||||
|
Language-specific guidance ({lang}):
|
||||||
|
- Output valid {lang} code.
|
||||||
|
- Prefer modern syntax and avoid prose unless comments are needed."""
|
||||||
|
|
||||||
|
|
||||||
|
def _language_guidance(language_id: str) -> str:
|
||||||
|
canonical = _canonical_language_id(language_id)
|
||||||
|
if canonical == "markdown":
|
||||||
|
return ""
|
||||||
|
guidance = _LANG_GUIDANCE.get(canonical)
|
||||||
|
if guidance:
|
||||||
|
return guidance
|
||||||
|
if canonical in _JS_LANGS:
|
||||||
|
return _JS_CODE.format(lang=canonical)
|
||||||
|
if canonical in _CODE_LANGS:
|
||||||
|
return _GENERIC_CODE.format(lang=canonical)
|
||||||
|
return _GENERIC_CODE.format(lang=canonical)
|
||||||
|
|
||||||
|
|
||||||
def build_inline_system_prompt(language_id: str = "markdown") -> str:
|
def build_inline_system_prompt(language_id: str = "markdown") -> str:
|
||||||
safe_language_id = _canonical_language_id(language_id)
|
safe_language_id = _canonical_language_id(language_id)
|
||||||
@@ -330,82 +331,103 @@ def build_inline_system_prompt(language_id: str = "markdown") -> str:
|
|||||||
|
|
||||||
Return only the insertion text that should be placed between PREFIX and SUFFIX.
|
Return only the insertion text that should be placed between PREFIX and SUFFIX.
|
||||||
|
|
||||||
Hard constraints you must follow:
|
CORE PRINCIPLE: Output insertion text only. No explanations, no meta labels, no wrapper quotes.
|
||||||
1) Output-only contract:
|
|
||||||
- Output insertion text only.
|
|
||||||
- No explanations, no meta labels, no wrapper quotes around the whole answer.
|
|
||||||
|
|
||||||
2) Strict math formatting (KaTeX):
|
PRIORITY 1: CONTEXT AWARENESS (Read these flags from user prompt)
|
||||||
- If you output any math expression, it must be strict KaTeX-compatible math.
|
- CURSOR_IN_FENCED_CODE_BLOCK: Are you inside a code fence?
|
||||||
- Every formula must be wrapped with either $...$ (inline) or $$...$$ (block).
|
- CURSOR_FENCE_LANGUAGE: What language is the current fence?
|
||||||
- Never output bare formulas without $ or $$ wrappers.
|
- PREFIX_ENDS_WITH_NEWLINE: Does prefix end with newline?
|
||||||
- Exception: If CURSOR_IN_FENCED_CODE_BLOCK=true and CURSOR_FENCE_LANGUAGE is latex/tex/katex,
|
- SUFFIX_STARTS_WITH_NEWLINE: Does suffix start with newline?
|
||||||
output raw LaTeX without $ or $$ wrappers.
|
- MERMAID_CONTEXT: Is this a Mermaid diagram context?
|
||||||
|
|
||||||
3) Strict code formatting:
|
PRIORITY 2: SPECIALIZED CONTENT RULES
|
||||||
- Read CURSOR_IN_FENCED_CODE_BLOCK from the user prompt.
|
|
||||||
- If CURSOR_IN_FENCED_CODE_BLOCK=true:
|
2.1 Code Block Handling:
|
||||||
- You are already inside a fenced code block.
|
If CURSOR_IN_FENCED_CODE_BLOCK=true:
|
||||||
- Never output triple backticks.
|
- You are inside a code fence
|
||||||
- Output code lines only.
|
- Output code lines ONLY (no triple backticks)
|
||||||
- If CURSOR_IN_FENCED_CODE_BLOCK=false:
|
- Use single \\n for code line separation
|
||||||
- Any code output must be in a fenced code block with a language tag:
|
|
||||||
|
If CURSOR_IN_FENCED_CODE_BLOCK=false and code needed:
|
||||||
|
- Wrap code in fenced block with language tag:
|
||||||
```{{language}}
|
```{{language}}
|
||||||
...
|
code here
|
||||||
```
|
```
|
||||||
- Do not output code snippets as inline backticks.
|
- Never use inline backticks for code snippets
|
||||||
- Choose the language tag from context (no default fallback tag instruction).
|
|
||||||
|
|
||||||
4) Mermaid-specific completion rules:
|
2.2 Math Formatting (KaTeX):
|
||||||
- Read CURSOR_FENCE_LANGUAGE and MERMAID_CONTEXT from the user prompt.
|
- Inline math: wrap with $...$
|
||||||
- If CURSOR_FENCE_LANGUAGE=mermaid:
|
- Block math: wrap with $$...$$
|
||||||
- Output Mermaid statements only.
|
- Never output bare formulas
|
||||||
- Never output triple backticks.
|
- Exception: inside latex/tex/katex fence, output raw LaTeX
|
||||||
- Never output prose explanations.
|
|
||||||
- If CURSOR_IN_FENCED_CODE_BLOCK=false and MERMAID_CONTEXT=true:
|
2.3 Mermaid Diagrams:
|
||||||
- Output a complete Mermaid fenced block:
|
If CURSOR_FENCE_LANGUAGE=mermaid:
|
||||||
|
- Output Mermaid syntax ONLY
|
||||||
|
- No backticks, no explanations
|
||||||
|
|
||||||
|
If MERMAID_CONTEXT=true and outside fence:
|
||||||
|
- Output complete fenced block:
|
||||||
```mermaid
|
```mermaid
|
||||||
...
|
diagram syntax
|
||||||
```
|
```
|
||||||
- Keep Mermaid syntax valid and concise.
|
|
||||||
- Never mix Mermaid code and explanatory narration in one output.
|
|
||||||
|
|
||||||
5) Boundary newline repair:
|
PRIORITY 3: MARKDOWN STRUCTURE
|
||||||
- Read PREFIX_ENDS_WITH_NEWLINE and SUFFIX_STARTS_WITH_NEWLINE from the user prompt.
|
|
||||||
- Carefully reason about whether OUTPUT should start or end with a newline.
|
|
||||||
- If PREFIX lacks a required boundary newline, add it at OUTPUT start.
|
|
||||||
- If SUFFIX lacks a required boundary newline, add it at OUTPUT end.
|
|
||||||
- Ensure PREFIX + OUTPUT + SUFFIX is structurally natural.
|
|
||||||
|
|
||||||
6) Context stitching:
|
3.1 Newline Semantics:
|
||||||
- Do not repeat text that already appears at the start of SUFFIX.
|
- Single \\n: soft break (same paragraph, renders as space or <br>)
|
||||||
- Preserve nearby language, tone, punctuation, indentation, and markdown structure.
|
- Double \\n\\n: hard break (new paragraph/block)
|
||||||
- Continue existing structures naturally (lists, tables, block quotes, headings).
|
- Use \\n\\n for: new paragraphs, before headings, starting lists/tables
|
||||||
|
- Use \\n for: continuation within blocks (list items, table cells)
|
||||||
|
- Exception: inside code blocks, use \\n freely for code lines
|
||||||
|
|
||||||
7) OCR safety:
|
3.2 Boundary Management:
|
||||||
- PREFIX may include hidden OCR metadata tags like <OCR:...>.
|
Check PREFIX_ENDS_WITH_NEWLINE and SUFFIX_STARTS_WITH_NEWLINE:
|
||||||
- Never output any OCR tag.
|
- If PREFIX lacks needed newline: start OUTPUT with \\n
|
||||||
- Never output OCR tag fragments such as <OCR:...>."""
|
- If SUFFIX lacks needed newline: end OUTPUT with \\n
|
||||||
|
- Common cases requiring leading \\n:
|
||||||
|
* Starting a list after "Steps:"
|
||||||
|
* Creating new paragraph after text
|
||||||
|
* Adding heading after paragraph
|
||||||
|
- Common cases requiring trailing \\n:
|
||||||
|
* Before new heading
|
||||||
|
* End of section
|
||||||
|
|
||||||
|
3.3 Context Stitching:
|
||||||
|
- Never repeat text from SUFFIX beginning
|
||||||
|
- Match PREFIX tone, style, indentation
|
||||||
|
- Continue structures: lists, tables, quotes, headings
|
||||||
|
|
||||||
|
PRIORITY 4: HIDDEN CONTEXT
|
||||||
|
- OCR metadata like <OCR:...> is hidden context
|
||||||
|
- Never copy OCR tags to output
|
||||||
|
- Use OCR content as semantic hint only
|
||||||
|
"""
|
||||||
|
|
||||||
if language_guidance:
|
if language_guidance:
|
||||||
system_prompt = f"{system_prompt.rstrip()}\n{language_guidance.strip()}"
|
system_prompt = f"{system_prompt.rstrip()}\\n{language_guidance.strip()}"
|
||||||
|
|
||||||
return system_prompt.strip()
|
return system_prompt.strip()
|
||||||
|
|
||||||
|
|
||||||
INLINE_EXAMPLES = """[EX01] Prose continuation
|
INLINE_EXAMPLES = """=== CATEGORY A: PROSE CONTINUATION ===
|
||||||
|
|
||||||
|
[EX01] Simple prose continuation
|
||||||
<PREFIX>The quick brown fox </PREFIX>
|
<PREFIX>The quick brown fox </PREFIX>
|
||||||
<SUFFIX>jumps over the lazy dog.</SUFFIX>
|
<SUFFIX>jumps over the lazy dog.</SUFFIX>
|
||||||
Expected OUTPUT:
|
Expected OUTPUT:
|
||||||
moved quietly and then
|
moved quietly and then
|
||||||
|
|
||||||
[EX02] Avoid repeating suffix beginning
|
[EX02] Avoid repeating suffix
|
||||||
<PREFIX>Our launch plan starts with </PREFIX>
|
<PREFIX>Our launch plan starts with </PREFIX>
|
||||||
<SUFFIX>phase one, followed by phase two.</SUFFIX>
|
<SUFFIX>phase one, followed by phase two.</SUFFIX>
|
||||||
Expected OUTPUT:
|
Expected OUTPUT:
|
||||||
careful internal testing before
|
careful internal testing before
|
||||||
|
WRONG: phase one starts with (repeats suffix)
|
||||||
|
|
||||||
[EX03] Continue markdown checklist
|
=== CATEGORY B: MARKDOWN STRUCTURES ===
|
||||||
|
|
||||||
|
[EX03] Continue checklist
|
||||||
<PREFIX>## TODO
|
<PREFIX>## TODO
|
||||||
- [ ] Buy milk
|
- [ ] Buy milk
|
||||||
- [ ] </PREFIX>
|
- [ ] </PREFIX>
|
||||||
@@ -413,41 +435,7 @@ careful internal testing before
|
|||||||
Expected OUTPUT:
|
Expected OUTPUT:
|
||||||
Write release notes and share draft with team
|
Write release notes and share draft with team
|
||||||
|
|
||||||
[EX04] Cursor outside code block, code must use fenced block
|
[EX04] Start list after header (PREFIX lacks newline)
|
||||||
CURSOR_IN_FENCED_CODE_BLOCK=false
|
|
||||||
<PREFIX>Parse this JSON payload in Python:</PREFIX>
|
|
||||||
<SUFFIX></SUFFIX>
|
|
||||||
Expected OUTPUT:
|
|
||||||
```python
|
|
||||||
import json
|
|
||||||
data = json.loads(payload)
|
|
||||||
```
|
|
||||||
|
|
||||||
[EX05] Cursor inside fenced code block, do not output fences
|
|
||||||
CURSOR_IN_FENCED_CODE_BLOCK=true
|
|
||||||
<PREFIX>```python
|
|
||||||
def add(a, b):
|
|
||||||
return </PREFIX>
|
|
||||||
<SUFFIX>
|
|
||||||
```</SUFFIX>
|
|
||||||
Expected OUTPUT:
|
|
||||||
a + b
|
|
||||||
|
|
||||||
[EX06] Inline math must use $...$
|
|
||||||
<PREFIX>The derivative of x^2 is </PREFIX>
|
|
||||||
<SUFFIX>.</SUFFIX>
|
|
||||||
Expected OUTPUT:
|
|
||||||
$2x$
|
|
||||||
|
|
||||||
[EX07] Block math must use $$...$$
|
|
||||||
<PREFIX>We can write the Gaussian integral as:</PREFIX>
|
|
||||||
<SUFFIX></SUFFIX>
|
|
||||||
Expected OUTPUT:
|
|
||||||
$$
|
|
||||||
\\int_{-\\infty}^{\\infty} e^{-x^2}\\,dx = \\sqrt{\\pi}
|
|
||||||
$$
|
|
||||||
|
|
||||||
[EX08] Prefix misses boundary newline; add newline at output start
|
|
||||||
PREFIX_ENDS_WITH_NEWLINE=false
|
PREFIX_ENDS_WITH_NEWLINE=false
|
||||||
<PREFIX>Deployment steps:</PREFIX>
|
<PREFIX>Deployment steps:</PREFIX>
|
||||||
<SUFFIX></SUFFIX>
|
<SUFFIX></SUFFIX>
|
||||||
@@ -456,21 +444,7 @@ Expected OUTPUT:
|
|||||||
- Build artifact
|
- Build artifact
|
||||||
- Deploy service
|
- Deploy service
|
||||||
|
|
||||||
[EX09] Suffix misses boundary newline; add newline at output end
|
[EX05] Continue table row
|
||||||
SUFFIX_STARTS_WITH_NEWLINE=false
|
|
||||||
<PREFIX>Summary paragraph complete.</PREFIX>
|
|
||||||
<SUFFIX>## Next Section</SUFFIX>
|
|
||||||
Expected OUTPUT:
|
|
||||||
|
|
||||||
|
|
||||||
[EX10] OCR metadata exists but must never be emitted
|
|
||||||
<PREFIX> <OCR:equation y = mx + b>
|
|
||||||
The relationship is </PREFIX>
|
|
||||||
<SUFFIX>.</SUFFIX>
|
|
||||||
Expected OUTPUT:
|
|
||||||
$y = mx + b$
|
|
||||||
|
|
||||||
[EX11] Continue markdown table with correct row shape
|
|
||||||
<PREFIX>| Name | Score |
|
<PREFIX>| Name | Score |
|
||||||
| --- | --- |
|
| --- | --- |
|
||||||
| Alice | 92 |
|
| Alice | 92 |
|
||||||
@@ -479,30 +453,91 @@ $y = mx + b$
|
|||||||
Expected OUTPUT:
|
Expected OUTPUT:
|
||||||
88 |
|
88 |
|
||||||
|
|
||||||
[EX12] Mixed text + math + code in one insertion
|
[EX06] Start new paragraph
|
||||||
CURSOR_IN_FENCED_CODE_BLOCK=false
|
<PREFIX>First paragraph ends.</PREFIX>
|
||||||
<PREFIX>Use the area formula and provide a tiny JS helper.</PREFIX>
|
|
||||||
<SUFFIX></SUFFIX>
|
<SUFFIX></SUFFIX>
|
||||||
Expected OUTPUT:
|
Expected OUTPUT:
|
||||||
The area is $A = \\pi r^2$.
|
|
||||||
|
|
||||||
```javascript
|
Second paragraph starts.
|
||||||
const area = (r) => Math.PI * r * r;
|
WRONG: Second paragraph starts. (missing leading \\n\\n)
|
||||||
|
|
||||||
|
[EX07] Add newline before heading
|
||||||
|
PREFIX_ENDS_WITH_NEWLINE=false
|
||||||
|
<PREFIX>End of previous section.</PREFIX>
|
||||||
|
<SUFFIX>## Next Heading</SUFFIX>
|
||||||
|
Expected OUTPUT:
|
||||||
|
|
||||||
|
WRONG: (would join with heading without separation)
|
||||||
|
|
||||||
|
=== CATEGORY C: CODE BLOCKS ===
|
||||||
|
|
||||||
|
[EX08] Outside fence: wrap code in fence
|
||||||
|
CURSOR_IN_FENCED_CODE_BLOCK=false
|
||||||
|
<PREFIX>Parse this JSON payload in Python:</PREFIX>
|
||||||
|
<SUFFIX></SUFFIX>
|
||||||
|
Expected OUTPUT:
|
||||||
|
```python
|
||||||
|
import json
|
||||||
|
data = json.loads(payload)
|
||||||
```
|
```
|
||||||
|
WRONG: import json\\ndata = json.loads(payload) (no fence)
|
||||||
|
|
||||||
[EX13] Cursor inside mermaid fence: no backticks, mermaid lines only
|
[EX09] Inside fence: output code only
|
||||||
CURSOR_IN_FENCED_CODE_BLOCK=true
|
CURSOR_IN_FENCED_CODE_BLOCK=true
|
||||||
|
<PREFIX>```python
|
||||||
|
def add(a, b):
|
||||||
|
return </PREFIX>
|
||||||
|
<SUFFIX>
|
||||||
|
```</SUFFIX>
|
||||||
|
Expected OUTPUT:
|
||||||
|
a + b
|
||||||
|
WRONG: ```python\\nreturn a + b\\n``` (duplicate fences)
|
||||||
|
|
||||||
|
[EX10] Code inside fence uses single newline
|
||||||
|
CURSOR_IN_FENCED_CODE_BLOCK=true
|
||||||
|
<PREFIX>```python
|
||||||
|
def hello():</PREFIX>
|
||||||
|
<SUFFIX>
|
||||||
|
```</SUFFIX>
|
||||||
|
Expected OUTPUT:
|
||||||
|
print("Hello")
|
||||||
|
return True
|
||||||
|
(Note: single \\n between code lines, no markdown rules)
|
||||||
|
|
||||||
|
=== CATEGORY D: MATH ===
|
||||||
|
|
||||||
|
[EX11] Inline math
|
||||||
|
<PREFIX>The derivative of x^2 is </PREFIX>
|
||||||
|
<SUFFIX>.</SUFFIX>
|
||||||
|
Expected OUTPUT:
|
||||||
|
$2x$
|
||||||
|
WRONG: 2x (bare formula)
|
||||||
|
|
||||||
|
[EX12] Block math
|
||||||
|
<PREFIX>We can write the Gaussian integral as:</PREFIX>
|
||||||
|
<SUFFIX></SUFFIX>
|
||||||
|
Expected OUTPUT:
|
||||||
|
$$
|
||||||
|
\\int_{-\\infty}^{\\infty} e^{-x^2}\\,dx = \\sqrt{\\pi}
|
||||||
|
$$
|
||||||
|
WRONG: \\int... (bare formula without $$)
|
||||||
|
|
||||||
|
=== CATEGORY E: MERMAID ===
|
||||||
|
|
||||||
|
[EX13] Inside mermaid fence
|
||||||
CURSOR_FENCE_LANGUAGE=mermaid
|
CURSOR_FENCE_LANGUAGE=mermaid
|
||||||
|
CURSOR_IN_FENCED_CODE_BLOCK=true
|
||||||
<PREFIX>```mermaid
|
<PREFIX>```mermaid
|
||||||
flowchart TD
|
flowchart TD
|
||||||
A[Start] --> </PREFIX>
|
A[Start] --> </PREFIX>
|
||||||
<SUFFIX>
|
<SUFFIX>
|
||||||
```</SUFFIX>
|
```</SUFFIX>
|
||||||
Expected OUTPUT:
|
Expected OUTPUT:
|
||||||
B{Valid?}
|
B{Valid?}
|
||||||
B -->|Yes| C[Done]
|
B -->|Yes| C[Done]
|
||||||
|
WRONG: ```mermaid\\nB{Valid?}... (duplicate fence)
|
||||||
|
|
||||||
[EX14] Mermaid context outside fence: return full mermaid block
|
[EX14] Outside fence with mermaid context
|
||||||
CURSOR_IN_FENCED_CODE_BLOCK=false
|
CURSOR_IN_FENCED_CODE_BLOCK=false
|
||||||
MERMAID_CONTEXT=true
|
MERMAID_CONTEXT=true
|
||||||
<PREFIX>Please provide a simple release pipeline diagram.</PREFIX>
|
<PREFIX>Please provide a simple release pipeline diagram.</PREFIX>
|
||||||
@@ -510,8 +545,18 @@ MERMAID_CONTEXT=true
|
|||||||
Expected OUTPUT:
|
Expected OUTPUT:
|
||||||
```mermaid
|
```mermaid
|
||||||
flowchart LR
|
flowchart LR
|
||||||
Build --> Test --> Deploy
|
Build --> Test --> Deploy
|
||||||
```"""
|
```
|
||||||
|
|
||||||
|
=== CATEGORY F: OCR METADATA ===
|
||||||
|
|
||||||
|
[EX15] Use OCR as context, never output
|
||||||
|
<PREFIX> <OCR:equation y = mx + b>
|
||||||
|
The relationship is </PREFIX>
|
||||||
|
<SUFFIX>.</SUFFIX>
|
||||||
|
Expected OUTPUT:
|
||||||
|
$y = mx + b$
|
||||||
|
WRONG: <OCR:equation y = mx + b> (OCR tag in output)"""
|
||||||
|
|
||||||
|
|
||||||
def build_completion_prompts(
|
def build_completion_prompts(
|
||||||
@@ -520,7 +565,7 @@ def build_completion_prompts(
|
|||||||
language_id: str = "markdown",
|
language_id: str = "markdown",
|
||||||
location: str = "",
|
location: str = "",
|
||||||
thinking_level: str = "low",
|
thinking_level: str = "low",
|
||||||
preferences: object = None,
|
preferences: UserPreferences | None = None,
|
||||||
) -> Tuple[str, str]:
|
) -> Tuple[str, str]:
|
||||||
safe_language_id = _canonical_language_id(language_id)
|
safe_language_id = _canonical_language_id(language_id)
|
||||||
recent_prefix, recent_suffix = _prepare_context(prefix, suffix)
|
recent_prefix, recent_suffix = _prepare_context(prefix, suffix)
|
||||||
@@ -551,35 +596,49 @@ def build_completion_prompts(
|
|||||||
preferences_instruction = f"\nUser Preferences:\n{preferences_instruction}"
|
preferences_instruction = f"\nUser Preferences:\n{preferences_instruction}"
|
||||||
|
|
||||||
user_prompt = f"""Current time: {current_time}{location_info}{preferences_instruction}
|
user_prompt = f"""Current time: {current_time}{location_info}{preferences_instruction}
|
||||||
Reasoning hint: {thinking_level}
|
Reasoning level: {thinking_level}
|
||||||
Editor language id: {safe_language_id}
|
Editor language: {safe_language_id}
|
||||||
|
|
||||||
Completion state flags:
|
=== STATE FLAGS ===
|
||||||
- CURSOR_IN_FENCED_CODE_BLOCK: {"true" if cursor_in_fenced_code_block else "false"}
|
- CURSOR_IN_FENCED_CODE_BLOCK: {"true" if cursor_in_fenced_code_block else "false"}
|
||||||
- CURSOR_FENCE_LANGUAGE: {cursor_fence_language}
|
- CURSOR_FENCE_LANGUAGE: {cursor_fence_language}
|
||||||
- MERMAID_CONTEXT: {"true" if mermaid_context else "false"}
|
- MERMAID_CONTEXT: {"true" if mermaid_context else "false"}
|
||||||
- PREFIX_ENDS_WITH_NEWLINE: {"true" if prefix_ends_with_newline else "false"}
|
- PREFIX_ENDS_WITH_NEWLINE: {"true" if prefix_ends_with_newline else "false"}
|
||||||
- SUFFIX_STARTS_WITH_NEWLINE: {"true" if suffix_starts_with_newline else "false"}
|
- SUFFIX_STARTS_WITH_NEWLINE: {"true" if suffix_starts_with_newline else "false"}
|
||||||
|
|
||||||
Task:
|
=== TASK ===
|
||||||
- Produce the best insertion text at the cursor between PREFIX and SUFFIX.
|
Produce the best insertion text between PREFIX and SUFFIX.
|
||||||
- Keep insertion meaningful and non-empty.
|
Requirements:
|
||||||
- Keep insertion concise unless structure requires more content.
|
- Non-empty and meaningful
|
||||||
|
- Concise unless structure needs more
|
||||||
|
- Follows markdown rules in system prompt
|
||||||
|
|
||||||
Context notes:
|
=== BOUNDARY DECISION GUIDE ===
|
||||||
- PREFIX may include OCR metadata after image markdown, e.g.  <OCR:description>.
|
|
||||||
- OCR metadata is hidden context and must never be copied into output.
|
|
||||||
- Preserve local style and formatting.
|
|
||||||
|
|
||||||
Decision policy:
|
Step 1: Check PREFIX_ENDS_WITH_NEWLINE
|
||||||
- Prioritize seamless join: PREFIX + OUTPUT + SUFFIX must read naturally.
|
If false, ask: "Does output need to start on a new line?"
|
||||||
- Do not repeat SUFFIX-leading text.
|
- YES if PREFIX ends with: ":", "steps:", "items:", heading text, or complete sentence before heading
|
||||||
- If uncertain, prefer a complete short phrase/sentence with clear meaning.
|
- If YES: start output with \\n
|
||||||
|
|
||||||
Comprehensive examples:
|
Step 2: Check SUFFIX_STARTS_WITH_NEWLINE
|
||||||
|
If false, ask: "Does output need to end with a newline?"
|
||||||
|
- YES if SUFFIX starts with: heading (##), new paragraph, or list marker
|
||||||
|
- If YES: end output with \\n
|
||||||
|
|
||||||
|
Step 3: Choose newline type
|
||||||
|
- Use \\n\\n for: new paragraphs, before headings, starting lists
|
||||||
|
- Use \\n for: continuing within blocks, list items, table cells
|
||||||
|
- Exception: inside code fences, use \\n freely
|
||||||
|
|
||||||
|
=== CONTEXT NOTES ===
|
||||||
|
- OCR metadata (e.g., <OCR:description>) is hidden context, never copy to output
|
||||||
|
- Match PREFIX tone, style, and indentation
|
||||||
|
- Do not repeat text from SUFFIX beginning
|
||||||
|
|
||||||
|
=== EXAMPLES BY CATEGORY ===
|
||||||
{INLINE_EXAMPLES}
|
{INLINE_EXAMPLES}
|
||||||
|
|
||||||
Now produce the insertion.
|
=== NOW COMPLETE THE TASK ===
|
||||||
|
|
||||||
<PREFIX>
|
<PREFIX>
|
||||||
{recent_prefix}
|
{recent_prefix}
|
||||||
@@ -601,7 +660,7 @@ def build_prompt(
|
|||||||
language_id: str = "markdown",
|
language_id: str = "markdown",
|
||||||
location: str = "",
|
location: str = "",
|
||||||
thinking_level: str = "low",
|
thinking_level: str = "low",
|
||||||
preferences: object = None,
|
preferences: UserPreferences | None = None,
|
||||||
) -> str:
|
) -> str:
|
||||||
"""
|
"""
|
||||||
Backward-compatible helper. Returns only the user prompt body.
|
Backward-compatible helper. Returns only the user prompt body.
|
||||||
|
|||||||
@@ -1,4 +1,4 @@
|
|||||||
fastapi
|
fastapi
|
||||||
uvicorn
|
uvicorn
|
||||||
ollama
|
ollama
|
||||||
pydantic
|
pydantic
|
||||||
@@ -10,3 +10,10 @@ python-docx
|
|||||||
python-pptx
|
python-pptx
|
||||||
openpyxl
|
openpyxl
|
||||||
pypdf
|
pypdf
|
||||||
|
|
||||||
|
# TTS and ASR dependencies
|
||||||
|
torch
|
||||||
|
transformers
|
||||||
|
soundfile
|
||||||
|
numpy
|
||||||
|
accelerate
|
||||||
|
|||||||
@@ -0,0 +1,256 @@
|
|||||||
|
# TTS and ASR API for macOS Silicon with HuggingFace transformers
|
||||||
|
import asyncio
|
||||||
|
import base64
|
||||||
|
import logging
|
||||||
|
import os
|
||||||
|
import platform
|
||||||
|
|
||||||
|
from fastapi import APIRouter, HTTPException, Security
|
||||||
|
from pydantic import BaseModel
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
router = APIRouter()
|
||||||
|
logger = logging.getLogger("tts_asr")
|
||||||
|
|
||||||
|
_tts_pipeline = None
|
||||||
|
_asr_pipeline = None
|
||||||
|
_device = None
|
||||||
|
|
||||||
|
|
||||||
|
def _get_device():
|
||||||
|
global _device
|
||||||
|
if _device is not None:
|
||||||
|
return _device
|
||||||
|
|
||||||
|
import torch
|
||||||
|
|
||||||
|
if platform.system() == "Darwin" and hasattr(torch.backends, "mps") and torch.backends.mps.is_available():
|
||||||
|
_device = "mps"
|
||||||
|
logger.info("[Device] 使用 MPS 加速")
|
||||||
|
elif torch.cuda.is_available():
|
||||||
|
_device = "cuda"
|
||||||
|
logger.info("[Device] 使用 CUDA 加速")
|
||||||
|
else:
|
||||||
|
_device = "cpu"
|
||||||
|
logger.info("[Device] 使用 CPU")
|
||||||
|
return _device
|
||||||
|
|
||||||
|
|
||||||
|
def _device_arg():
|
||||||
|
device = _get_device()
|
||||||
|
if device == "cuda":
|
||||||
|
return "cuda:0"
|
||||||
|
return device
|
||||||
|
|
||||||
|
|
||||||
|
def _get_tts_pipeline():
|
||||||
|
global _tts_pipeline
|
||||||
|
if _tts_pipeline is not None:
|
||||||
|
return _tts_pipeline
|
||||||
|
|
||||||
|
import torch
|
||||||
|
from transformers import pipeline
|
||||||
|
|
||||||
|
logger.info("[TTS] 加载 Kokoro-82M 模型...")
|
||||||
|
_tts_pipeline = pipeline(
|
||||||
|
"text-to-speech",
|
||||||
|
model="hexgrad/Kokoro-82M",
|
||||||
|
trust_remote_code=True,
|
||||||
|
device=_device_arg(),
|
||||||
|
torch_dtype=torch.float16 if _get_device() != "cpu" else torch.float32,
|
||||||
|
)
|
||||||
|
logger.info("[TTS] Kokoro-82M 模型加载完成")
|
||||||
|
return _tts_pipeline
|
||||||
|
|
||||||
|
|
||||||
|
def _get_asr_pipeline():
|
||||||
|
global _asr_pipeline
|
||||||
|
if _asr_pipeline is not None:
|
||||||
|
return _asr_pipeline
|
||||||
|
|
||||||
|
import torch
|
||||||
|
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor, pipeline
|
||||||
|
|
||||||
|
logger.info("[ASR] 加载 Whisper large-v3-turbo 模型...")
|
||||||
|
model_id = "openai/whisper-large-v3-turbo"
|
||||||
|
model = AutoModelForSpeechSeq2Seq.from_pretrained(
|
||||||
|
model_id,
|
||||||
|
torch_dtype=torch.float16 if _get_device() != "cpu" else torch.float32,
|
||||||
|
low_cpu_mem_usage=True,
|
||||||
|
use_safetensors=True,
|
||||||
|
)
|
||||||
|
processor = AutoProcessor.from_pretrained(model_id)
|
||||||
|
_asr_pipeline = pipeline(
|
||||||
|
"automatic-speech-recognition",
|
||||||
|
model=model,
|
||||||
|
tokenizer=processor.tokenizer,
|
||||||
|
feature_extractor=processor.feature_extractor,
|
||||||
|
torch_dtype=torch.float16 if _get_device() != "cpu" else torch.float32,
|
||||||
|
device=_device_arg(),
|
||||||
|
)
|
||||||
|
logger.info("[ASR] Whisper large-v3-turbo 模型加载完成")
|
||||||
|
return _asr_pipeline
|
||||||
|
|
||||||
|
|
||||||
|
def _save_audio_to_wav(audio_data: bytes, sample_rate: int = 16000) -> str:
|
||||||
|
import tempfile
|
||||||
|
import wave
|
||||||
|
|
||||||
|
with tempfile.NamedTemporaryFile(suffix=".wav", delete=False, mode="wb") as tmp:
|
||||||
|
with wave.open(tmp.name, "wb") as wf:
|
||||||
|
wf.setnchannels(1)
|
||||||
|
wf.setsampwidth(2)
|
||||||
|
wf.setframerate(sample_rate)
|
||||||
|
wf.writeframes(audio_data)
|
||||||
|
return tmp.name
|
||||||
|
|
||||||
|
|
||||||
|
def _tts_sync(text: str, voice: str = "af_bella", rate: float = 1.0) -> tuple[bytes, int]:
|
||||||
|
tts = _get_tts_pipeline()
|
||||||
|
result = tts(text, voice=voice)
|
||||||
|
audio = None
|
||||||
|
sample_rate = 24000
|
||||||
|
if isinstance(result, dict):
|
||||||
|
audio = result.get("audio")
|
||||||
|
sample_rate = int(result.get("sampling_rate", sample_rate))
|
||||||
|
elif isinstance(result, (list, tuple)) and result:
|
||||||
|
audio = result[0]
|
||||||
|
|
||||||
|
if audio is None:
|
||||||
|
raise RuntimeError("Kokoro 未返回音频数据")
|
||||||
|
|
||||||
|
if hasattr(audio, "cpu"):
|
||||||
|
audio = audio.cpu().numpy()
|
||||||
|
|
||||||
|
duration_ms = int(len(audio) * 1000 / sample_rate)
|
||||||
|
|
||||||
|
if audio.dtype != np.int16:
|
||||||
|
audio = (audio * 32767).astype(np.int16)
|
||||||
|
|
||||||
|
import tempfile
|
||||||
|
import wave
|
||||||
|
|
||||||
|
with tempfile.NamedTemporaryFile(suffix=".wav", delete=False) as tmp:
|
||||||
|
output_path = tmp.name
|
||||||
|
try:
|
||||||
|
with wave.open(output_path, "wb") as wf:
|
||||||
|
wf.setnchannels(1)
|
||||||
|
wf.setsampwidth(2)
|
||||||
|
wf.setframerate(sample_rate)
|
||||||
|
wf.writeframes(audio.tobytes())
|
||||||
|
with open(output_path, "rb") as f:
|
||||||
|
return f.read(), duration_ms
|
||||||
|
finally:
|
||||||
|
if os.path.exists(output_path):
|
||||||
|
os.unlink(output_path)
|
||||||
|
|
||||||
|
|
||||||
|
async def _text_to_speech(text: str, voice: str = "af_bella", rate: float = 1.0) -> tuple[bytes, int]:
|
||||||
|
return await asyncio.to_thread(_tts_sync, text, voice, rate)
|
||||||
|
|
||||||
|
|
||||||
|
def _asr_sync(audio_data: bytes, language: str = "zh") -> str:
|
||||||
|
import soundfile as sf
|
||||||
|
|
||||||
|
asr = _get_asr_pipeline()
|
||||||
|
audio_path = _save_audio_to_wav(audio_data)
|
||||||
|
try:
|
||||||
|
audio_array, sample_rate = sf.read(audio_path)
|
||||||
|
result = asr(
|
||||||
|
audio_array,
|
||||||
|
sampling_rate=sample_rate,
|
||||||
|
generate_kwargs={"language": language, "task": "transcribe"},
|
||||||
|
)
|
||||||
|
if isinstance(result, dict):
|
||||||
|
return result.get("text", "").strip()
|
||||||
|
return str(result).strip()
|
||||||
|
finally:
|
||||||
|
if os.path.exists(audio_path):
|
||||||
|
os.unlink(audio_path)
|
||||||
|
|
||||||
|
|
||||||
|
async def _speech_to_text(audio_data: bytes, language: str = "zh") -> str:
|
||||||
|
return await asyncio.to_thread(_asr_sync, audio_data, language)
|
||||||
|
|
||||||
|
|
||||||
|
class TTSRequest(BaseModel):
|
||||||
|
text: str
|
||||||
|
voice: str = "af_bella"
|
||||||
|
rate: float = 1.0
|
||||||
|
format: str = "wav"
|
||||||
|
|
||||||
|
|
||||||
|
class TTSResponse(BaseModel):
|
||||||
|
audio_base64: str
|
||||||
|
format: str
|
||||||
|
duration_ms: int
|
||||||
|
|
||||||
|
|
||||||
|
class ASRRequest(BaseModel):
|
||||||
|
audio_base64: str
|
||||||
|
language: str = "zh-CN"
|
||||||
|
|
||||||
|
|
||||||
|
class ASRResponse(BaseModel):
|
||||||
|
text: str
|
||||||
|
language: str
|
||||||
|
|
||||||
|
|
||||||
|
def get_api_key(api_key: str):
|
||||||
|
import main
|
||||||
|
|
||||||
|
API_KEY = main.API_KEY
|
||||||
|
if api_key != API_KEY:
|
||||||
|
raise HTTPException(status_code=403, detail="API Key 无效")
|
||||||
|
return api_key
|
||||||
|
|
||||||
|
|
||||||
|
@router.post("/tts", response_model=TTSResponse)
|
||||||
|
async def text_to_speech(req: TTSRequest, api_key: str = Security(get_api_key)):
|
||||||
|
request_id = str(hash(req.text))[:8]
|
||||||
|
try:
|
||||||
|
logger.info("[TTS][%s] text_chars=%d voice=%s format=%s", request_id, len(req.text), req.voice, req.format)
|
||||||
|
audio_data, duration_ms = await _text_to_speech(req.text, req.voice, req.rate)
|
||||||
|
if req.format.lower() == "mp3":
|
||||||
|
import subprocess
|
||||||
|
import tempfile
|
||||||
|
|
||||||
|
with tempfile.NamedTemporaryFile(suffix=".wav", delete=False) as tmp_in:
|
||||||
|
tmp_in.write(audio_data)
|
||||||
|
input_path = tmp_in.name
|
||||||
|
with tempfile.NamedTemporaryFile(suffix=".mp3", delete=False) as tmp_out:
|
||||||
|
output_path = tmp_out.name
|
||||||
|
try:
|
||||||
|
cmd = ["ffmpeg", "-i", input_path, "-acodec", "libmp3lame", "-ab", "128k", output_path]
|
||||||
|
result = await asyncio.to_thread(lambda: subprocess.run(cmd, capture_output=True, text=True, timeout=30))
|
||||||
|
if result.returncode != 0:
|
||||||
|
raise RuntimeError(f"MP3 转换失败: {result.stderr}")
|
||||||
|
with open(output_path, "rb") as f:
|
||||||
|
audio_data = f.read()
|
||||||
|
finally:
|
||||||
|
for path in [input_path, output_path]:
|
||||||
|
if os.path.exists(path):
|
||||||
|
os.unlink(path)
|
||||||
|
logger.info("[TTS][%s] success duration_ms=%d", request_id, duration_ms)
|
||||||
|
return TTSResponse(audio_base64=base64.b64encode(audio_data).decode(), format=req.format, duration_ms=duration_ms)
|
||||||
|
except Exception as e:
|
||||||
|
logger.exception("[TTS] failed: %s", e)
|
||||||
|
raise HTTPException(status_code=500, detail=str(e))
|
||||||
|
|
||||||
|
|
||||||
|
@router.post("/asr", response_model=ASRResponse)
|
||||||
|
async def speech_to_text(req: ASRRequest, api_key: str = Security(get_api_key)):
|
||||||
|
request_id = str(hash(req.audio_base64))[:8]
|
||||||
|
try:
|
||||||
|
logger.info("[ASR][%s] audio_base64_chars=%d language=%s", request_id, len(req.audio_base64), req.language)
|
||||||
|
audio_data = base64.b64decode(req.audio_base64)
|
||||||
|
text = await _speech_to_text(audio_data, req.language[:2])
|
||||||
|
logger.info("[ASR][%s] success text_chars=%d", request_id, len(text))
|
||||||
|
return ASRResponse(text=text, language=req.language)
|
||||||
|
except Exception as e:
|
||||||
|
logger.exception("[ASR] failed: %s", e)
|
||||||
|
raise HTTPException(status_code=500, detail=str(e))
|
||||||
|
|
||||||
|
|
||||||
|
def register_tts_asr_routes(app):
|
||||||
|
app.include_router(router, prefix="/v1/tts-asr")
|
||||||
Generated
+1543
File diff suppressed because it is too large
Load Diff
@@ -11,13 +11,17 @@
|
|||||||
"check": "npm run build"
|
"check": "npm run build"
|
||||||
},
|
},
|
||||||
"dependencies": {
|
"dependencies": {
|
||||||
|
"@blocknote/xl-docx-exporter": "^0.47.3",
|
||||||
"@milkdown/core": "^7.18.0",
|
"@milkdown/core": "^7.18.0",
|
||||||
"@milkdown/crepe": "^7.18.0",
|
"@milkdown/crepe": "^7.18.0",
|
||||||
"@milkdown/kit": "^7.18.0",
|
"@milkdown/kit": "^7.18.0",
|
||||||
"@milkdown/theme-nord": "^7.18.0",
|
"@milkdown/theme-nord": "^7.18.0",
|
||||||
"@milkdown/vue": "^7.18.0",
|
"@milkdown/vue": "^7.18.0",
|
||||||
"docx": "^9.6.0",
|
"docx": "^9.6.0",
|
||||||
|
"docx-preview": "^0.3.7",
|
||||||
|
"docx2pdf-converter": "^2.1.1",
|
||||||
"html2pdf.js": "^0.14.0",
|
"html2pdf.js": "^0.14.0",
|
||||||
|
"jspdf": "^4.2.1",
|
||||||
"katex": "^0.16.9",
|
"katex": "^0.16.9",
|
||||||
"markdown-it": "^13.0.0",
|
"markdown-it": "^13.0.0",
|
||||||
"markdown-it-math": "^3.0.2",
|
"markdown-it-math": "^3.0.2",
|
||||||
|
|||||||
+207
-146
@@ -1,250 +1,311 @@
|
|||||||
<template>
|
<template>
|
||||||
<div class="doc-block-crepe" :class="{ collapsed: isCollapsed }">
|
<section class="doc-card" :class="{ 'is-collapsed': collapsedState }">
|
||||||
<div class="doc-header">
|
<header class="doc-card__header">
|
||||||
<div class="doc-icon">
|
<div class="doc-card__badge">{{ typeLabel }}</div>
|
||||||
<svg v-if="docType === 'pdf'" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2">
|
<div class="doc-card__meta">
|
||||||
<path d="M14 2H6a2 2 0 0 0-2 2v16a2 2 0 0 0 2 2h12a2 2 0 0 0 2-2V8z"/>
|
<div class="doc-card__name">{{ docName }}</div>
|
||||||
<polyline points="14 2 14 8 20 8"/>
|
<div class="doc-card__time">{{ displayTime }}</div>
|
||||||
<path d="M9 15v-2h6v2"/>
|
|
||||||
<path d="M12 13v4"/>
|
|
||||||
</svg>
|
|
||||||
<svg v-else-if="docType === 'doc'" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2">
|
|
||||||
<path d="M14 2H6a2 2 0 0 0-2 2v16a2 2 0 0 0 2 2h12a2 2 0 0 0 2-2V8z"/>
|
|
||||||
<polyline points="14 2 14 8 20 8"/>
|
|
||||||
<path d="M16 13H8"/>
|
|
||||||
<path d="M16 17H8"/>
|
|
||||||
<path d="M10 9H8"/>
|
|
||||||
</svg>
|
|
||||||
<svg v-else-if="docType === 'ppt'" width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2">
|
|
||||||
<rect x="2" y="3" width="20" height="14" rx="2"/>
|
|
||||||
<path d="M8 21h8"/>
|
|
||||||
<path d="M12 17v4"/>
|
|
||||||
</svg>
|
|
||||||
<svg v-else width="18" height="18" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2">
|
|
||||||
<path d="M14 2H6a2 2 0 0 0-2 2v16a2 2 0 0 0 2 2h12a2 2 0 0 0 2-2V8z"/>
|
|
||||||
<polyline points="14 2 14 8 20 8"/>
|
|
||||||
<path d="M16 13H8"/>
|
|
||||||
<path d="M16 17H8"/>
|
|
||||||
</svg>
|
|
||||||
</div>
|
</div>
|
||||||
<div class="doc-name">{{ docName }}</div>
|
<div class="doc-card__actions">
|
||||||
<div class="doc-actions">
|
<button type="button" class="doc-card__btn" :title="collapsedState ? '展开文件' : '折叠文件'" @click="toggleCollapse">
|
||||||
<button @click="downloadDoc" class="action-btn" title="下载文档">
|
<svg v-if="collapsedState" width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2">
|
||||||
<svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2">
|
|
||||||
<path d="M21 15v4a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2v-4"/>
|
|
||||||
<polyline points="7 10 12 15 17 10"/>
|
|
||||||
<line x1="12" y1="15" x2="12" y2="3"/>
|
|
||||||
</svg>
|
|
||||||
</button>
|
|
||||||
<button @click="toggleCollapse" class="action-btn collapse-btn" :title="isCollapsed ? '展开' : '折叠'">
|
|
||||||
<svg v-if="isCollapsed" width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2">
|
|
||||||
<polyline points="9 18 15 12 9 6"/>
|
<polyline points="9 18 15 12 9 6"/>
|
||||||
</svg>
|
</svg>
|
||||||
<svg v-else width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2">
|
<svg v-else width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2">
|
||||||
<polyline points="6 9 12 15 18 9"/>
|
<polyline points="6 9 12 15 18 9"/>
|
||||||
</svg>
|
</svg>
|
||||||
</button>
|
</button>
|
||||||
|
<button type="button" class="doc-card__btn doc-card__btn--danger" title="删除文件" @click="props.onDelete?.()">
|
||||||
|
<svg width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2">
|
||||||
|
<path d="M3 6h18"/>
|
||||||
|
<path d="M8 6V4h8v2"/>
|
||||||
|
<path d="M19 6l-1 14H6L5 6"/>
|
||||||
|
<path d="M10 11v6"/>
|
||||||
|
<path d="M14 11v6"/>
|
||||||
|
</svg>
|
||||||
|
</button>
|
||||||
</div>
|
</div>
|
||||||
|
</header>
|
||||||
|
<div v-show="!collapsedState" class="doc-card__body">
|
||||||
|
<div ref="editorRoot" class="doc-card__editor"></div>
|
||||||
</div>
|
</div>
|
||||||
<div class="doc-editor" v-show="!isCollapsed">
|
</section>
|
||||||
<div ref="editorRoot" class="inner-crepe"></div>
|
|
||||||
</div>
|
|
||||||
</div>
|
|
||||||
</template>
|
</template>
|
||||||
|
|
||||||
<script setup>
|
<script setup>
|
||||||
import { ref, onMounted, onUnmounted, watch } from 'vue'
|
import { computed, onMounted, onUnmounted, ref, watch } from 'vue'
|
||||||
|
import { replaceAll } from '@milkdown/kit/utils'
|
||||||
import { Crepe } from '@milkdown/crepe'
|
import { Crepe } from '@milkdown/crepe'
|
||||||
import { editorViewCtx, serializerCtx } from '@milkdown/kit/core'
|
import { editorViewCtx } from '@milkdown/kit/core'
|
||||||
import { copilotPlugin, copilotConfigCtx, setCopilotEnabled } from '../plugins/copilotPlugin'
|
import { copilotPlugin, copilotConfigCtx, copilotGhostMark, setCopilotEnabled, clearGhostSuggestion } from '../plugins/copilotPlugin'
|
||||||
import { fetchSuggestion } from '../utils/api.js'
|
import { fetchSuggestion } from '../utils/api.js'
|
||||||
|
|
||||||
const props = defineProps({
|
const props = defineProps({
|
||||||
docType: { type: String, default: 'text' },
|
docType: { type: String, default: 'txt' },
|
||||||
docName: { type: String, default: 'document.txt' },
|
docName: { type: String, default: 'document.txt' },
|
||||||
uploadTime: { type: String, default: '' },
|
uploadTime: { type: String, default: '' },
|
||||||
initialContent: { type: String, default: '' }
|
content: { type: String, default: '' },
|
||||||
|
collapsed: { type: Boolean, default: false },
|
||||||
|
resolveSuggestionRequest: { type: Function, default: null },
|
||||||
|
onUpdateContent: { type: Function, default: null },
|
||||||
|
onUpdateCollapsed: { type: Function, default: null },
|
||||||
|
onDelete: { type: Function, default: null },
|
||||||
})
|
})
|
||||||
|
|
||||||
const emit = defineEmits(['update:content', 'delete'])
|
|
||||||
|
|
||||||
const editorRoot = ref(null)
|
const editorRoot = ref(null)
|
||||||
const isCollapsed = ref(false)
|
const collapsedState = ref(Boolean(props.collapsed))
|
||||||
|
const currentContent = ref(props.content || '')
|
||||||
let crepe = null
|
let crepe = null
|
||||||
let internalChangeTimer = null
|
let syncTimer = null
|
||||||
|
let syncingExternal = false
|
||||||
|
|
||||||
|
const typeLabel = computed(() => {
|
||||||
|
if (props.docType === 'docx') return 'DOCX'
|
||||||
|
if (props.docType === 'pptx') return 'PPTX'
|
||||||
|
if (props.docType === 'pdf') return 'PDF'
|
||||||
|
return 'TXT'
|
||||||
|
})
|
||||||
|
|
||||||
|
const displayTime = computed(() => {
|
||||||
|
if (!props.uploadTime) return '刚上传'
|
||||||
|
const date = new Date(props.uploadTime)
|
||||||
|
if (Number.isNaN(date.getTime())) return '刚上传'
|
||||||
|
return date.toLocaleString('zh-CN', { hour12: false })
|
||||||
|
})
|
||||||
|
|
||||||
const toggleCollapse = () => {
|
const toggleCollapse = () => {
|
||||||
isCollapsed.value = !isCollapsed.value
|
collapsedState.value = !collapsedState.value
|
||||||
}
|
props.onUpdateCollapsed?.(collapsedState.value)
|
||||||
|
|
||||||
const downloadDoc = () => {
|
|
||||||
if (!crepe) return
|
|
||||||
crepe.getMarkdown().then(markdown => {
|
|
||||||
const blob = new Blob([markdown], { type: 'text/plain;charset=utf-8' })
|
|
||||||
const url = URL.createObjectURL(blob)
|
|
||||||
const a = document.createElement('a')
|
|
||||||
a.href = url
|
|
||||||
a.download = props.docName
|
|
||||||
document.body.appendChild(a)
|
|
||||||
a.click()
|
|
||||||
a.remove()
|
|
||||||
URL.revokeObjectURL(url)
|
|
||||||
})
|
|
||||||
}
|
}
|
||||||
|
|
||||||
const syncContent = () => {
|
const syncContent = () => {
|
||||||
if (!crepe) return
|
if (!crepe) return
|
||||||
if (internalChangeTimer) clearTimeout(internalChangeTimer)
|
if (syncTimer) clearTimeout(syncTimer)
|
||||||
internalChangeTimer = setTimeout(async () => {
|
syncTimer = setTimeout(async () => {
|
||||||
|
if (!crepe || syncingExternal) return
|
||||||
const markdown = await crepe.getMarkdown()
|
const markdown = await crepe.getMarkdown()
|
||||||
emit('update:content', markdown)
|
currentContent.value = markdown
|
||||||
|
props.onUpdateContent?.(markdown)
|
||||||
}, 120)
|
}, 120)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
const syncExternalContent = async (nextValue) => {
|
||||||
|
const value = nextValue || ''
|
||||||
|
if (!crepe) {
|
||||||
|
currentContent.value = value
|
||||||
|
return
|
||||||
|
}
|
||||||
|
if (value === currentContent.value) return
|
||||||
|
syncingExternal = true
|
||||||
|
try {
|
||||||
|
crepe.editor.action(replaceAll(value))
|
||||||
|
currentContent.value = value
|
||||||
|
} finally {
|
||||||
|
syncingExternal = false
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
watch(() => props.content, (nextValue) => {
|
||||||
|
void syncExternalContent(nextValue)
|
||||||
|
})
|
||||||
|
|
||||||
|
watch(() => props.collapsed, (nextValue) => {
|
||||||
|
collapsedState.value = Boolean(nextValue)
|
||||||
|
})
|
||||||
|
|
||||||
onMounted(async () => {
|
onMounted(async () => {
|
||||||
if (!editorRoot.value) return
|
if (!editorRoot.value) return
|
||||||
crepe = new Crepe({
|
crepe = new Crepe({
|
||||||
root: editorRoot.value,
|
root: editorRoot.value,
|
||||||
defaultValue: props.initialContent || '',
|
defaultValue: props.content || '',
|
||||||
features: {
|
features: {
|
||||||
[Crepe.Feature.Latex]: true,
|
[Crepe.Feature.Latex]: true,
|
||||||
[Crepe.Feature.ImageBlock]: true,
|
[Crepe.Feature.ImageBlock]: true,
|
||||||
[Crepe.Feature.Table]: true,
|
[Crepe.Feature.Table]: true,
|
||||||
[Crepe.Feature.ListCheck]: true,
|
[Crepe.Feature.ListCheck]: true,
|
||||||
},
|
},
|
||||||
config: { showLineNumber: false }
|
config: {
|
||||||
|
showLineNumber: false,
|
||||||
|
},
|
||||||
})
|
})
|
||||||
|
|
||||||
crepe.editor.config(ctx => {
|
crepe.editor.config((ctx) => {
|
||||||
ctx.set(copilotConfigCtx.key, {
|
ctx.set(copilotConfigCtx.key, {
|
||||||
fetchSuggestion,
|
fetchSuggestion: async (prefix, suffix, languageId, signal) => {
|
||||||
debounceMs: 1000
|
const payload = props.resolveSuggestionRequest
|
||||||
|
? await props.resolveSuggestionRequest({ prefix, suffix, languageId })
|
||||||
|
: { prefix, suffix, languageId, blocked: false }
|
||||||
|
if (payload?.blocked) return ''
|
||||||
|
return fetchSuggestion(payload?.prefix ?? prefix, payload?.suffix ?? suffix, payload?.languageId ?? languageId, signal)
|
||||||
|
},
|
||||||
|
debounceMs: 900,
|
||||||
})
|
})
|
||||||
})
|
})
|
||||||
|
|
||||||
|
crepe.editor.use(copilotConfigCtx)
|
||||||
|
crepe.editor.use(copilotGhostMark)
|
||||||
crepe.editor.use(copilotPlugin)
|
crepe.editor.use(copilotPlugin)
|
||||||
await crepe.create()
|
await crepe.create()
|
||||||
|
|
||||||
crepe.on(listener => {
|
crepe.on((listener) => {
|
||||||
listener.updated(() => {
|
listener.updated(() => {
|
||||||
syncContent()
|
syncContent()
|
||||||
})
|
})
|
||||||
})
|
})
|
||||||
|
|
||||||
crepe.editor.action(ctx => {
|
crepe.editor.action((ctx) => {
|
||||||
const view = ctx.get(editorViewCtx)
|
const view = ctx.get(editorViewCtx)
|
||||||
setCopilotEnabled(view, true)
|
setCopilotEnabled(view, true)
|
||||||
})
|
})
|
||||||
})
|
})
|
||||||
|
|
||||||
watch(() => props.initialContent, (newVal) => {
|
|
||||||
if (crepe && newVal !== undefined) {
|
|
||||||
crepe.editor.action(ctx => {
|
|
||||||
const view = ctx.get(editorViewCtx)
|
|
||||||
const currentPos = view.state.selection.from
|
|
||||||
view.dispatch(view.state.tr.insertText(newVal))
|
|
||||||
})
|
|
||||||
}
|
|
||||||
})
|
|
||||||
|
|
||||||
onUnmounted(() => {
|
onUnmounted(() => {
|
||||||
if (internalChangeTimer) clearTimeout(internalChangeTimer)
|
if (syncTimer) {
|
||||||
|
clearTimeout(syncTimer)
|
||||||
|
syncTimer = null
|
||||||
|
}
|
||||||
if (crepe) {
|
if (crepe) {
|
||||||
|
crepe.editor.action((ctx) => {
|
||||||
|
const view = ctx.get(editorViewCtx)
|
||||||
|
clearGhostSuggestion(view)
|
||||||
|
})
|
||||||
crepe.destroy()
|
crepe.destroy()
|
||||||
crepe = null
|
crepe = null
|
||||||
}
|
}
|
||||||
})
|
})
|
||||||
|
|
||||||
defineExpose({
|
|
||||||
getContent: () => crepe ? crepe.getMarkdown() : Promise.resolve(''),
|
|
||||||
getEditor: () => crepe
|
|
||||||
})
|
|
||||||
</script>
|
</script>
|
||||||
|
|
||||||
<style scoped>
|
<style scoped>
|
||||||
.doc-block-crepe {
|
.doc-card {
|
||||||
margin: 12px 0;
|
width: 100%;
|
||||||
border-radius: 8px;
|
max-width: 100%;
|
||||||
|
margin: 8px 0;
|
||||||
|
border-radius: 12px;
|
||||||
|
border: 1px solid rgba(59, 130, 246, 0.12);
|
||||||
|
background: rgba(255, 255, 255, 0.78);
|
||||||
|
box-shadow: 0 2px 8px rgba(59, 130, 246, 0.06), 0 1px 3px rgba(0, 0, 0, 0.04);
|
||||||
overflow: hidden;
|
overflow: hidden;
|
||||||
background: var(--crepe-color-surface-low);
|
backdrop-filter: blur(10px);
|
||||||
border: 1px solid var(--panel-border);
|
position: relative;
|
||||||
}
|
}
|
||||||
|
|
||||||
.doc-block-crepe.collapsed .doc-editor {
|
.doc-card__header {
|
||||||
display: none;
|
display: grid;
|
||||||
}
|
grid-template-columns: auto minmax(0, 1fr) auto;
|
||||||
|
|
||||||
.doc-header {
|
|
||||||
display: flex;
|
|
||||||
align-items: center;
|
|
||||||
padding: 10px 12px;
|
|
||||||
background: var(--crepe-color-surface);
|
|
||||||
border-bottom: 1px solid var(--panel-border);
|
|
||||||
gap: 10px;
|
gap: 10px;
|
||||||
}
|
|
||||||
|
|
||||||
.doc-icon {
|
|
||||||
display: flex;
|
|
||||||
align-items: center;
|
align-items: center;
|
||||||
justify-content: center;
|
padding: 8px 12px;
|
||||||
color: var(--crepe-color-primary);
|
border-bottom: 1px solid rgba(59, 130, 246, 0.1);
|
||||||
|
background: rgba(255, 255, 255, 0.6);
|
||||||
}
|
}
|
||||||
|
|
||||||
.doc-name {
|
.doc-card__badge {
|
||||||
flex: 1;
|
min-width: 48px;
|
||||||
font-size: 14px;
|
padding: 4px 10px;
|
||||||
font-weight: 500;
|
border-radius: 999px;
|
||||||
color: var(--crepe-color-on-surface);
|
background: linear-gradient(135deg, #3b82f6 0%, #60a5fa 100%);
|
||||||
|
color: #fff;
|
||||||
|
font-size: 10px;
|
||||||
|
font-weight: 600;
|
||||||
|
letter-spacing: 0.08em;
|
||||||
|
text-align: center;
|
||||||
|
box-shadow: 0 2px 6px rgba(59, 130, 246, 0.2);
|
||||||
|
}
|
||||||
|
|
||||||
|
.doc-card__meta {
|
||||||
|
min-width: 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
.doc-card__name {
|
||||||
|
color: #1e293b;
|
||||||
|
font-size: 13px;
|
||||||
|
font-weight: 600;
|
||||||
|
white-space: nowrap;
|
||||||
overflow: hidden;
|
overflow: hidden;
|
||||||
text-overflow: ellipsis;
|
text-overflow: ellipsis;
|
||||||
white-space: nowrap;
|
|
||||||
}
|
}
|
||||||
|
|
||||||
.doc-actions {
|
.doc-card__time {
|
||||||
|
margin-top: 2px;
|
||||||
|
color: #64748b;
|
||||||
|
font-size: 10px;
|
||||||
|
}
|
||||||
|
|
||||||
|
.doc-card__actions {
|
||||||
display: flex;
|
display: flex;
|
||||||
align-items: center;
|
|
||||||
gap: 4px;
|
gap: 4px;
|
||||||
}
|
}
|
||||||
|
|
||||||
.action-btn {
|
.doc-card__btn {
|
||||||
|
width: 26px;
|
||||||
|
height: 26px;
|
||||||
|
border: 1px solid rgba(59, 130, 246, 0.12);
|
||||||
|
border-radius: 8px;
|
||||||
|
background: rgba(255, 255, 255, 0.5);
|
||||||
|
color: #64748b;
|
||||||
display: flex;
|
display: flex;
|
||||||
align-items: center;
|
align-items: center;
|
||||||
justify-content: center;
|
justify-content: center;
|
||||||
width: 28px;
|
|
||||||
height: 28px;
|
|
||||||
padding: 0;
|
|
||||||
border: none;
|
|
||||||
background: transparent;
|
|
||||||
color: var(--crepe-color-on-surface-variant);
|
|
||||||
cursor: pointer;
|
cursor: pointer;
|
||||||
border-radius: 4px;
|
transition: all 0.15s ease;
|
||||||
opacity: 0.7;
|
|
||||||
}
|
}
|
||||||
|
|
||||||
.action-btn:hover {
|
.doc-card__btn:hover {
|
||||||
background: var(--crepe-color-hover);
|
background: rgba(59, 130, 246, 0.1);
|
||||||
opacity: 1;
|
border-color: rgba(59, 130, 246, 0.25);
|
||||||
|
color: #3b82f6;
|
||||||
}
|
}
|
||||||
|
|
||||||
.doc-editor {
|
.doc-card__btn--danger:hover {
|
||||||
padding: 8px;
|
background: rgba(239, 68, 68, 0.1);
|
||||||
background: var(--crepe-color-surface-low);
|
border-color: rgba(239, 68, 68, 0.2);
|
||||||
min-height: 120px;
|
color: #ef4444;
|
||||||
max-height: 400px;
|
|
||||||
overflow-y: auto;
|
|
||||||
}
|
}
|
||||||
|
|
||||||
.inner-crepe {
|
.doc-card__body {
|
||||||
width: 100%;
|
padding: 8px 10px;
|
||||||
height: 100%;
|
background: rgba(248, 250, 252, 0.5);
|
||||||
}
|
}
|
||||||
|
|
||||||
.inner-crepe :deep(.milkdown) {
|
.doc-card__editor {
|
||||||
|
min-height: 48px;
|
||||||
|
border-radius: 8px;
|
||||||
|
border: 1px solid rgba(59, 130, 246, 0.08);
|
||||||
|
background: rgba(255, 255, 255, 0.6);
|
||||||
|
overflow: hidden;
|
||||||
|
}
|
||||||
|
|
||||||
|
.doc-card__editor :deep(.milkdown) {
|
||||||
background: transparent !important;
|
background: transparent !important;
|
||||||
}
|
}
|
||||||
|
|
||||||
.inner-crepe :deep(.ProseMirror) {
|
.doc-card__editor :deep(.milkdown__main),
|
||||||
|
.doc-card__editor :deep(.milkdown__editor) {
|
||||||
|
margin: 0 !important;
|
||||||
|
padding: 0 !important;
|
||||||
|
}
|
||||||
|
|
||||||
|
.doc-card__editor :deep(.ProseMirror) {
|
||||||
min-height: 80px;
|
min-height: 80px;
|
||||||
padding: 8px !important;
|
padding: 10px 12px 12px !important;
|
||||||
|
font-size: 13px !important;
|
||||||
|
line-height: 1.6;
|
||||||
|
}
|
||||||
|
|
||||||
|
.doc-card__editor :deep(.ProseMirror > *:last-child) {
|
||||||
|
margin-bottom: 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
.doc-card__editor :deep(.ProseMirror p:first-child) {
|
||||||
|
margin-top: 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
.doc-card__editor :deep(.milkdown__toolbar),
|
||||||
|
.doc-card__editor :deep(.milkdown__menu),
|
||||||
|
.doc-card__editor :deep(.milkdown__statusbar),
|
||||||
|
.doc-card__editor :deep(.milkdown-slate-toolbar),
|
||||||
|
.doc-card__editor :deep(.milkdown-bubble-menu) {
|
||||||
|
display: none !important;
|
||||||
}
|
}
|
||||||
</style>
|
</style>
|
||||||
|
|
||||||
|
|||||||
@@ -107,78 +107,76 @@ const downloadDoc = () => {
|
|||||||
|
|
||||||
<style scoped>
|
<style scoped>
|
||||||
.doc-block {
|
.doc-block {
|
||||||
margin: 12px 0;
|
margin: 8px 0;
|
||||||
border-radius: 8px;
|
border-radius: 10px;
|
||||||
overflow: hidden;
|
overflow: hidden;
|
||||||
background: var(--crepe-color-surface-low);
|
background: rgba(255, 255, 255, 0.72);
|
||||||
border: 1px solid var(--panel-border);
|
backdrop-filter: blur(12px);
|
||||||
|
border: 1px solid rgba(59, 130, 246, 0.15);
|
||||||
|
box-shadow: 0 2px 8px rgba(59, 130, 246, 0.06), 0 1px 2px rgba(0, 0, 0, 0.04);
|
||||||
}
|
}
|
||||||
|
|
||||||
.doc-block.collapsed .doc-content {
|
.doc-block.collapsed .doc-content {
|
||||||
display: none;
|
display: none;
|
||||||
}
|
}
|
||||||
|
|
||||||
/* 深色条 */
|
|
||||||
.doc-header {
|
.doc-header {
|
||||||
display: flex;
|
display: flex;
|
||||||
align-items: center;
|
align-items: center;
|
||||||
padding: 10px 12px;
|
padding: 6px 10px;
|
||||||
background: var(--crepe-color-surface);
|
background: rgba(255, 255, 255, 0.85);
|
||||||
border-bottom: 1px solid var(--panel-border);
|
border-bottom: 1px solid rgba(59, 130, 246, 0.12);
|
||||||
gap: 10px;
|
gap: 8px;
|
||||||
}
|
}
|
||||||
|
|
||||||
/* 文件类型icon */
|
|
||||||
.doc-icon {
|
.doc-icon {
|
||||||
display: flex;
|
display: flex;
|
||||||
align-items: center;
|
align-items: center;
|
||||||
justify-content: center;
|
justify-content: center;
|
||||||
color: var(--crepe-color-primary);
|
color: #3b82f6;
|
||||||
}
|
}
|
||||||
|
|
||||||
/* 文件名 */
|
|
||||||
.doc-name {
|
.doc-name {
|
||||||
flex: 1;
|
flex: 1;
|
||||||
font-size: 14px;
|
font-size: 13px;
|
||||||
font-weight: 500;
|
font-weight: 500;
|
||||||
color: var(--crepe-color-on-surface);
|
color: #1e293b;
|
||||||
overflow: hidden;
|
overflow: hidden;
|
||||||
text-overflow: ellipsis;
|
text-overflow: ellipsis;
|
||||||
white-space: nowrap;
|
white-space: nowrap;
|
||||||
}
|
}
|
||||||
|
|
||||||
/* 操作按钮 */
|
|
||||||
.doc-actions {
|
.doc-actions {
|
||||||
display: flex;
|
display: flex;
|
||||||
align-items: center;
|
align-items: center;
|
||||||
gap: 4px;
|
gap: 2px;
|
||||||
}
|
}
|
||||||
|
|
||||||
.action-btn {
|
.action-btn {
|
||||||
display: flex;
|
display: flex;
|
||||||
align-items: center;
|
align-items: center;
|
||||||
justify-content: center;
|
justify-content: center;
|
||||||
width: 28px;
|
width: 24px;
|
||||||
height: 28px;
|
height: 24px;
|
||||||
padding: 0;
|
padding: 0;
|
||||||
border: none;
|
border: none;
|
||||||
background: transparent;
|
background: transparent;
|
||||||
color: var(--crepe-color-on-surface-variant);
|
color: #64748b;
|
||||||
cursor: pointer;
|
cursor: pointer;
|
||||||
border-radius: 4px;
|
border-radius: 6px;
|
||||||
opacity: 0.7;
|
opacity: 0.75;
|
||||||
}
|
}
|
||||||
|
|
||||||
.action-btn:hover {
|
.action-btn:hover {
|
||||||
background: var(--crepe-color-hover);
|
background: rgba(59, 130, 246, 0.1);
|
||||||
|
color: #3b82f6;
|
||||||
opacity: 1;
|
opacity: 1;
|
||||||
}
|
}
|
||||||
|
|
||||||
/* 浅色块:文档内容 */
|
|
||||||
.doc-content {
|
.doc-content {
|
||||||
padding: 12px;
|
padding: 8px 10px;
|
||||||
background: var(--crepe-color-surface-low);
|
background: rgba(248, 250, 252, 0.6);
|
||||||
max-height: 400px;
|
max-height: 240px;
|
||||||
overflow-y: auto;
|
overflow-y: auto;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -186,9 +184,9 @@ const downloadDoc = () => {
|
|||||||
margin: 0;
|
margin: 0;
|
||||||
padding: 0;
|
padding: 0;
|
||||||
font-family: 'SF Mono', 'Monaco', 'Inconsolata', 'Fira Mono', monospace;
|
font-family: 'SF Mono', 'Monaco', 'Inconsolata', 'Fira Mono', monospace;
|
||||||
font-size: 13px;
|
font-size: 12px;
|
||||||
line-height: 1.6;
|
line-height: 1.5;
|
||||||
color: var(--crepe-color-on-surface);
|
color: #334155;
|
||||||
white-space: pre-wrap;
|
white-space: pre-wrap;
|
||||||
word-break: break-word;
|
word-break: break-word;
|
||||||
}
|
}
|
||||||
|
|||||||
+623
-455
File diff suppressed because it is too large
Load Diff
@@ -743,7 +743,12 @@ export function interruptCopilot(view: EditorView): void {
|
|||||||
}
|
}
|
||||||
|
|
||||||
export function checkSizeLimit(view: EditorView): { size: number; overLimit: boolean } {
|
export function checkSizeLimit(view: EditorView): { size: number; overLimit: boolean } {
|
||||||
const size = view.state.doc.content.size
|
let size = view.state.doc.content.size
|
||||||
|
view.state.doc.descendants((node) => {
|
||||||
|
if (node.type.name === 'doc_block' && node.attrs.content) {
|
||||||
|
size += String(node.attrs.content).length
|
||||||
|
}
|
||||||
|
})
|
||||||
return { size, overLimit: size > SIZE_LIMIT }
|
return { size, overLimit: size > SIZE_LIMIT }
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,249 @@
|
|||||||
|
import { createApp, reactive } from 'vue'
|
||||||
|
import { serializerCtx } from '@milkdown/kit/core'
|
||||||
|
import { $node, $remark, $view } from '@milkdown/kit/utils'
|
||||||
|
import type { Node as ProseNode, Schema } from '@milkdown/prose/model'
|
||||||
|
import type { EditorView, NodeView } from '@milkdown/prose/view'
|
||||||
|
import DocBlockCrepe from '../components/DocBlockCrepe.vue'
|
||||||
|
import {
|
||||||
|
DOC_BLOCK_FENCE_LANG,
|
||||||
|
DOC_BLOCK_NODE_TYPE,
|
||||||
|
DOC_CONTEXT_LIMIT,
|
||||||
|
buildLegacyDocBlock,
|
||||||
|
buildDocContextFence,
|
||||||
|
normalizeDocType,
|
||||||
|
parseLegacyDocBlock,
|
||||||
|
parseDocBlockValue,
|
||||||
|
stripDocBlockMarkdown,
|
||||||
|
} from '../utils/docBlock.js'
|
||||||
|
|
||||||
|
function serializeRangeToMarkdown(
|
||||||
|
doc: ProseNode,
|
||||||
|
from: number,
|
||||||
|
to: number,
|
||||||
|
schema: Schema,
|
||||||
|
serializer: (content: ProseNode) => string
|
||||||
|
): string {
|
||||||
|
if (from >= to) return ''
|
||||||
|
const slice = doc.slice(from, to)
|
||||||
|
if (slice.content.size <= 0) return ''
|
||||||
|
const sliceDoc = schema.topNodeType.createAndFill(undefined, slice.content)
|
||||||
|
return sliceDoc ? serializer(sliceDoc) : doc.textBetween(from, to, '\n', '\n')
|
||||||
|
}
|
||||||
|
|
||||||
|
function buildDocContext(doc: ProseNode, excludePos?: number) {
|
||||||
|
const blocks: string[] = []
|
||||||
|
doc.descendants((node, pos) => {
|
||||||
|
if (node.type.name !== DOC_BLOCK_NODE_TYPE) return true
|
||||||
|
if (excludePos !== undefined && pos === excludePos) return false
|
||||||
|
blocks.push(
|
||||||
|
buildDocContextFence({
|
||||||
|
docType: node.attrs.docType,
|
||||||
|
content: node.attrs.content,
|
||||||
|
})
|
||||||
|
)
|
||||||
|
return false
|
||||||
|
})
|
||||||
|
return blocks.join('\n\n')
|
||||||
|
}
|
||||||
|
|
||||||
|
class DocBlockNodeView implements NodeView {
|
||||||
|
node: ProseNode
|
||||||
|
view: EditorView
|
||||||
|
getPos: () => number | undefined
|
||||||
|
dom: HTMLElement
|
||||||
|
app: ReturnType<typeof createApp> | null = null
|
||||||
|
props: Record<string, any>
|
||||||
|
serializer: (content: ProseNode) => string
|
||||||
|
|
||||||
|
constructor(node: ProseNode, view: EditorView, getPos: () => number | undefined, serializer: (content: ProseNode) => string) {
|
||||||
|
this.node = node
|
||||||
|
this.view = view
|
||||||
|
this.getPos = getPos
|
||||||
|
this.serializer = serializer
|
||||||
|
this.dom = document.createElement('div')
|
||||||
|
this.dom.className = 'doc-block-node-view'
|
||||||
|
this.props = reactive({
|
||||||
|
docType: node.attrs.docType,
|
||||||
|
docName: node.attrs.docName,
|
||||||
|
uploadTime: node.attrs.uploadTime,
|
||||||
|
content: node.attrs.content,
|
||||||
|
collapsed: node.attrs.collapsed,
|
||||||
|
onUpdateContent: (content: string) => this.updateAttrs({ content }),
|
||||||
|
onUpdateCollapsed: (collapsed: boolean) => this.updateAttrs({ collapsed }),
|
||||||
|
onDelete: () => this.deleteNode(),
|
||||||
|
resolveSuggestionRequest: (payload: { prefix: string; suffix: string; languageId: string }) => this.resolveSuggestionRequest(payload),
|
||||||
|
})
|
||||||
|
this.mount()
|
||||||
|
}
|
||||||
|
|
||||||
|
mount() {
|
||||||
|
this.app = createApp(DocBlockCrepe, this.props)
|
||||||
|
this.app.mount(this.dom)
|
||||||
|
}
|
||||||
|
|
||||||
|
getPosValue() {
|
||||||
|
const pos = this.getPos()
|
||||||
|
return typeof pos === 'number' ? pos : undefined
|
||||||
|
}
|
||||||
|
|
||||||
|
updateAttrs(patch: Record<string, any>) {
|
||||||
|
const pos = this.getPosValue()
|
||||||
|
if (pos === undefined) return
|
||||||
|
const nextAttrs = { ...this.node.attrs, ...patch }
|
||||||
|
this.view.dispatch(this.view.state.tr.setNodeMarkup(pos, undefined, nextAttrs))
|
||||||
|
}
|
||||||
|
|
||||||
|
deleteNode() {
|
||||||
|
const pos = this.getPosValue()
|
||||||
|
if (pos === undefined) return
|
||||||
|
const tr = this.view.state.tr.delete(pos, pos + this.node.nodeSize).scrollIntoView()
|
||||||
|
this.view.dispatch(tr)
|
||||||
|
this.view.focus()
|
||||||
|
}
|
||||||
|
|
||||||
|
resolveSuggestionRequest(payload: { prefix: string; suffix: string; languageId: string }) {
|
||||||
|
const pos = this.getPosValue()
|
||||||
|
if (pos === undefined) return payload
|
||||||
|
const doc = this.view.state.doc
|
||||||
|
const schema = this.view.state.schema
|
||||||
|
const before = stripDocBlockMarkdown(serializeRangeToMarkdown(doc, 0, pos, schema, this.serializer))
|
||||||
|
const after = stripDocBlockMarkdown(serializeRangeToMarkdown(doc, pos + this.node.nodeSize, doc.content.size, schema, this.serializer))
|
||||||
|
const docContext = buildDocContext(doc, pos)
|
||||||
|
const mergedPrefix = [docContext, before, payload.prefix].filter(Boolean).join('\n\n')
|
||||||
|
const mergedSuffix = [payload.suffix, after].filter(Boolean).join('\n\n')
|
||||||
|
if (mergedPrefix.length + mergedSuffix.length > DOC_CONTEXT_LIMIT) {
|
||||||
|
return {
|
||||||
|
prefix: mergedPrefix.slice(0, DOC_CONTEXT_LIMIT),
|
||||||
|
suffix: '',
|
||||||
|
languageId: payload.languageId,
|
||||||
|
blocked: true,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return {
|
||||||
|
prefix: mergedPrefix,
|
||||||
|
suffix: mergedSuffix,
|
||||||
|
languageId: payload.languageId,
|
||||||
|
blocked: false,
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
update(node: ProseNode) {
|
||||||
|
if (node.type !== this.node.type) return false
|
||||||
|
this.node = node
|
||||||
|
this.props.docType = node.attrs.docType
|
||||||
|
this.props.docName = node.attrs.docName
|
||||||
|
this.props.uploadTime = node.attrs.uploadTime
|
||||||
|
this.props.content = node.attrs.content
|
||||||
|
this.props.collapsed = node.attrs.collapsed
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
|
||||||
|
stopEvent(event: Event) {
|
||||||
|
const target = event.target as Node | null
|
||||||
|
return Boolean(target && this.dom.contains(target))
|
||||||
|
}
|
||||||
|
|
||||||
|
ignoreMutation() {
|
||||||
|
return true
|
||||||
|
}
|
||||||
|
|
||||||
|
destroy() {
|
||||||
|
this.app?.unmount()
|
||||||
|
this.app = null
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function visitChildren(node: any, visitor: (child: any) => any) {
|
||||||
|
if (!node || !Array.isArray(node.children)) return
|
||||||
|
node.children = node.children.map((child: any) => {
|
||||||
|
const next = visitor(child)
|
||||||
|
if (next && next !== child) return next
|
||||||
|
visitChildren(child, visitor)
|
||||||
|
return child
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
export const docBlockRemark = $remark('docBlockRemark', () => () => {
|
||||||
|
return (tree: any) => {
|
||||||
|
visitChildren(tree, (node) => {
|
||||||
|
if (node?.type === 'code' && node.lang === DOC_BLOCK_FENCE_LANG) {
|
||||||
|
return {
|
||||||
|
type: 'docBlock',
|
||||||
|
value: String(node.value || ''),
|
||||||
|
sourceType: 'code',
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if (node?.type === 'html' && typeof node.value === 'string' && node.value.includes('<doc_type=')) {
|
||||||
|
return {
|
||||||
|
type: 'docBlock',
|
||||||
|
value: String(node.value || ''),
|
||||||
|
sourceType: 'html',
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return node
|
||||||
|
})
|
||||||
|
}
|
||||||
|
})
|
||||||
|
|
||||||
|
export const docBlockNode = $node(DOC_BLOCK_NODE_TYPE, () => ({
|
||||||
|
group: 'block',
|
||||||
|
atom: true,
|
||||||
|
isolating: true,
|
||||||
|
selectable: true,
|
||||||
|
draggable: false,
|
||||||
|
marks: '',
|
||||||
|
attrs: {
|
||||||
|
docType: { default: 'txt' },
|
||||||
|
docName: { default: 'document.txt' },
|
||||||
|
uploadTime: { default: '' },
|
||||||
|
content: { default: '' },
|
||||||
|
collapsed: { default: false },
|
||||||
|
},
|
||||||
|
parseDOM: [
|
||||||
|
{
|
||||||
|
tag: 'div[data-doc-block="true"]',
|
||||||
|
getAttrs: (dom) => ({
|
||||||
|
docType: normalizeDocType((dom as HTMLElement).getAttribute('data-doc-type') || ''),
|
||||||
|
docName: (dom as HTMLElement).getAttribute('data-doc-name') || 'document.txt',
|
||||||
|
uploadTime: (dom as HTMLElement).getAttribute('data-doc-upload-time') || '',
|
||||||
|
collapsed: ((dom as HTMLElement).getAttribute('data-doc-collapsed') || '') === 'true',
|
||||||
|
content: '',
|
||||||
|
}),
|
||||||
|
},
|
||||||
|
],
|
||||||
|
toDOM: (node) => [
|
||||||
|
'div',
|
||||||
|
{
|
||||||
|
'data-doc-block': 'true',
|
||||||
|
'data-doc-type': node.attrs.docType,
|
||||||
|
'data-doc-name': node.attrs.docName,
|
||||||
|
'data-doc-upload-time': node.attrs.uploadTime,
|
||||||
|
'data-doc-collapsed': String(Boolean(node.attrs.collapsed)),
|
||||||
|
},
|
||||||
|
],
|
||||||
|
parseMarkdown: {
|
||||||
|
match: (node) => node.type === 'docBlock',
|
||||||
|
runner: (state, node, type) => {
|
||||||
|
const attrs = node.sourceType === 'code'
|
||||||
|
? parseDocBlockValue(String(node.value || ''))
|
||||||
|
: parseLegacyDocBlock(String(node.value || ''))
|
||||||
|
if (!attrs) return
|
||||||
|
state.addNode(type, attrs)
|
||||||
|
},
|
||||||
|
},
|
||||||
|
toMarkdown: {
|
||||||
|
match: (node) => node.type.name === DOC_BLOCK_NODE_TYPE,
|
||||||
|
runner: (state, node) => {
|
||||||
|
state.addNode('html', undefined, buildLegacyDocBlock(node.attrs))
|
||||||
|
},
|
||||||
|
},
|
||||||
|
}))
|
||||||
|
|
||||||
|
export const docBlockView = $view(docBlockNode, (ctx) => {
|
||||||
|
const serializer = ctx.get(serializerCtx)
|
||||||
|
return (node, view, getPos) => new DocBlockNodeView(node, view, getPos, serializer)
|
||||||
|
})
|
||||||
|
|
||||||
|
export function buildDocContextFromDoc(doc: ProseNode, excludePos?: number) {
|
||||||
|
return buildDocContext(doc, excludePos)
|
||||||
|
}
|
||||||
+4
-4
@@ -273,7 +273,7 @@ body {
|
|||||||
overflow: auto;
|
overflow: auto;
|
||||||
border-radius: 8px;
|
border-radius: 8px;
|
||||||
padding: 8px;
|
padding: 8px;
|
||||||
background: color-mix(in srgb, var(--crepe-color-background, #fff) 88%, transparent);
|
background: rgba(255, 255, 255, 0.88);
|
||||||
}
|
}
|
||||||
|
|
||||||
.mermaid-inner::-webkit-scrollbar {
|
.mermaid-inner::-webkit-scrollbar {
|
||||||
@@ -315,7 +315,7 @@ body {
|
|||||||
.mermaid-error {
|
.mermaid-error {
|
||||||
padding: 12px 16px;
|
padding: 12px 16px;
|
||||||
margin: 0;
|
margin: 0;
|
||||||
background: color-mix(in srgb, var(--danger-text, #dc2626) 8%, transparent);
|
background: rgba(220, 38, 38, 0.08);
|
||||||
border: 1px solid var(--danger-text, #dc2626);
|
border: 1px solid var(--danger-text, #dc2626);
|
||||||
border-radius: 6px;
|
border-radius: 6px;
|
||||||
color: var(--danger-text, #dc2626);
|
color: var(--danger-text, #dc2626);
|
||||||
@@ -331,12 +331,12 @@ body {
|
|||||||
|
|
||||||
:root[data-theme='dark'] .milkdown .cm-editor,
|
:root[data-theme='dark'] .milkdown .cm-editor,
|
||||||
:root[data-theme='dark'] .milkdown .cm-scroller {
|
:root[data-theme='dark'] .milkdown .cm-scroller {
|
||||||
background-color: color-mix(in srgb, var(--crepe-color-surface-low) 86%, transparent);
|
background-color: rgba(237, 237, 237, 0.86);
|
||||||
color: var(--crepe-color-on-surface);
|
color: var(--crepe-color-on-surface);
|
||||||
}
|
}
|
||||||
|
|
||||||
:root[data-theme='dark'] .milkdown .cm-gutters {
|
:root[data-theme='dark'] .milkdown .cm-gutters {
|
||||||
background-color: color-mix(in srgb, var(--crepe-color-surface-low) 86%, transparent);
|
background-color: rgba(237, 237, 237, 0.86);
|
||||||
color: var(--crepe-color-on-surface-variant);
|
color: var(--crepe-color-on-surface-variant);
|
||||||
border-right-color: var(--panel-border);
|
border-right-color: var(--panel-border);
|
||||||
}
|
}
|
||||||
|
|||||||
+5
-33
@@ -1,4 +1,4 @@
|
|||||||
import { API_URL } from './config.js'
|
import { API_URL, API_KEY } from './config.js'
|
||||||
import { useSettingsStore } from '../stores/settings'
|
import { useSettingsStore } from '../stores/settings'
|
||||||
|
|
||||||
function generateRequestId() {
|
function generateRequestId() {
|
||||||
@@ -30,6 +30,7 @@ async function sendCancelRequest(cancelUrl, requestId, reason) {
|
|||||||
method: 'POST',
|
method: 'POST',
|
||||||
headers: {
|
headers: {
|
||||||
'Content-Type': 'application/json',
|
'Content-Type': 'application/json',
|
||||||
|
'X-API-Key': API_KEY,
|
||||||
},
|
},
|
||||||
body: JSON.stringify({
|
body: JSON.stringify({
|
||||||
request_id: requestId,
|
request_id: requestId,
|
||||||
@@ -73,6 +74,7 @@ export async function fetchSuggestion(prefix, suffix, languageId, signal, apiUrl
|
|||||||
const headers = {
|
const headers = {
|
||||||
'Content-Type': 'application/json',
|
'Content-Type': 'application/json',
|
||||||
'X-Request-Id': requestId,
|
'X-Request-Id': requestId,
|
||||||
|
'X-API-Key': API_KEY,
|
||||||
}
|
}
|
||||||
|
|
||||||
const body = {
|
const body = {
|
||||||
@@ -100,38 +102,8 @@ export async function fetchSuggestion(prefix, suffix, languageId, signal, apiUrl
|
|||||||
throw new Error(`HTTP ${res.status}: ${errorText}`)
|
throw new Error(`HTTP ${res.status}: ${errorText}`)
|
||||||
}
|
}
|
||||||
|
|
||||||
const reader = res.body?.getReader()
|
const data = await res.json()
|
||||||
if (!reader) {
|
return data.content || ''
|
||||||
throw new Error('No reader available')
|
|
||||||
}
|
|
||||||
|
|
||||||
let text = ''
|
|
||||||
let buffer = ''
|
|
||||||
while (true) {
|
|
||||||
const { done, value } = await reader.read()
|
|
||||||
if (done) break
|
|
||||||
buffer += new TextDecoder().decode(value)
|
|
||||||
|
|
||||||
const lines = buffer.split('\n')
|
|
||||||
buffer = lines.pop() || ''
|
|
||||||
|
|
||||||
for (const line of lines) {
|
|
||||||
if (!line.startsWith('data: ')) continue
|
|
||||||
const jsonStr = line.slice(6).trim()
|
|
||||||
if (!jsonStr) continue
|
|
||||||
try {
|
|
||||||
const data = JSON.parse(jsonStr)
|
|
||||||
if (data.content) {
|
|
||||||
text += data.content
|
|
||||||
}
|
|
||||||
if (data.done || data.error) break
|
|
||||||
} catch (e) {
|
|
||||||
// skip invalid lines
|
|
||||||
}
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
return text
|
|
||||||
} catch (e) {
|
} catch (e) {
|
||||||
if (e.name === 'AbortError') {
|
if (e.name === 'AbortError') {
|
||||||
// ignore abort
|
// ignore abort
|
||||||
|
|||||||
@@ -5,3 +5,5 @@ const API_BASE_URL = import.meta.env.VITE_API_BASE_URL || 'https://api.imageteac
|
|||||||
export const API_URL = import.meta.env.VITE_API_URL || `${API_BASE_URL}/v1/completions`
|
export const API_URL = import.meta.env.VITE_API_URL || `${API_BASE_URL}/v1/completions`
|
||||||
export const OCR_URL = import.meta.env.VITE_OCR_URL || `${API_BASE_URL}/v1/ocr`
|
export const OCR_URL = import.meta.env.VITE_OCR_URL || `${API_BASE_URL}/v1/ocr`
|
||||||
export const CONVERT_URL = import.meta.env.VITE_CONVERT_URL || `${API_BASE_URL}/v1/convert`
|
export const CONVERT_URL = import.meta.env.VITE_CONVERT_URL || `${API_BASE_URL}/v1/convert`
|
||||||
|
export const EXPORT_PDF_URL = import.meta.env.VITE_EXPORT_PDF_URL || '/v1/export/pdf'
|
||||||
|
export const API_KEY = import.meta.env.VITE_API_KEY || 'your-secret-key-here'
|
||||||
|
|||||||
@@ -23,6 +23,7 @@ export async function convertFileToMarkdown(file) {
|
|||||||
method: 'POST',
|
method: 'POST',
|
||||||
headers: {
|
headers: {
|
||||||
'Content-Type': 'application/json',
|
'Content-Type': 'application/json',
|
||||||
|
'X-API-Key': 'your-secret-key-here',
|
||||||
},
|
},
|
||||||
body: JSON.stringify({
|
body: JSON.stringify({
|
||||||
file: base64,
|
file: base64,
|
||||||
|
|||||||
@@ -0,0 +1,194 @@
|
|||||||
|
export const DOC_BLOCK_NODE_TYPE = 'doc_block'
|
||||||
|
export const DOC_BLOCK_FENCE_LANG = 'llm-file'
|
||||||
|
export const DOC_CONTEXT_LIMIT = 32 * 1024
|
||||||
|
|
||||||
|
const IMAGE_MD_RE = /!\[[^\]]*]\([^)]+\)/g
|
||||||
|
const IMAGE_HTML_RE = /<img\b[^>]*>/gi
|
||||||
|
const HEADER_SEPARATOR = '\n---\n'
|
||||||
|
|
||||||
|
export function normalizeDocType(value = '') {
|
||||||
|
const lower = String(value || '').trim().toLowerCase()
|
||||||
|
if (lower === 'txt' || lower === 'text' || lower === 'plain') return 'txt'
|
||||||
|
if (lower === 'json') return 'json'
|
||||||
|
if (lower === 'toml') return 'toml'
|
||||||
|
if (lower === 'yaml' || lower === 'yml') return 'yaml'
|
||||||
|
if (lower === 'doc' || lower === 'docx' || lower === 'word') return 'docx'
|
||||||
|
if (lower === 'ppt' || lower === 'pptx' || lower === 'powerpoint') return 'pptx'
|
||||||
|
if (lower === 'pdf') return 'pdf'
|
||||||
|
return 'txt'
|
||||||
|
}
|
||||||
|
|
||||||
|
export function getDocTypeFromFilename(name = '') {
|
||||||
|
const lower = String(name || '').toLowerCase()
|
||||||
|
if (lower.endsWith('.docx')) return 'docx'
|
||||||
|
if (lower.endsWith('.pptx')) return 'pptx'
|
||||||
|
if (lower.endsWith('.pdf')) return 'pdf'
|
||||||
|
if (lower.endsWith('.json')) return 'json'
|
||||||
|
if (lower.endsWith('.toml')) return 'toml'
|
||||||
|
if (lower.endsWith('.yaml') || lower.endsWith('.yml')) return 'yaml'
|
||||||
|
return 'txt'
|
||||||
|
}
|
||||||
|
|
||||||
|
export function isSupportedDocFile(file) {
|
||||||
|
if (!file) return false
|
||||||
|
const name = String(file.name || '').toLowerCase()
|
||||||
|
const type = String(file.type || '').toLowerCase()
|
||||||
|
return (
|
||||||
|
name.endsWith('.txt') ||
|
||||||
|
name.endsWith('.json') ||
|
||||||
|
name.endsWith('.toml') ||
|
||||||
|
name.endsWith('.yaml') ||
|
||||||
|
name.endsWith('.yml') ||
|
||||||
|
name.endsWith('.docx') ||
|
||||||
|
name.endsWith('.pptx') ||
|
||||||
|
name.endsWith('.pdf') ||
|
||||||
|
type === 'text/plain' ||
|
||||||
|
type === 'application/json' ||
|
||||||
|
type === 'text/yaml' ||
|
||||||
|
type === 'text/x-yaml' ||
|
||||||
|
type === 'application/x-yaml' ||
|
||||||
|
type === 'application/vnd.openxmlformats-officedocument.wordprocessingml.document' ||
|
||||||
|
type === 'application/vnd.openxmlformats-officedocument.presentationml.presentation' ||
|
||||||
|
type === 'application/pdf'
|
||||||
|
)
|
||||||
|
}
|
||||||
|
|
||||||
|
export function sanitizeDocContent(markdown = '') {
|
||||||
|
return String(markdown || '')
|
||||||
|
.replace(/\r\n?/g, '\n')
|
||||||
|
.replace(IMAGE_MD_RE, '')
|
||||||
|
.replace(IMAGE_HTML_RE, '')
|
||||||
|
.replace(/\n{3,}/g, '\n\n')
|
||||||
|
.trim()
|
||||||
|
}
|
||||||
|
|
||||||
|
function quoteMeta(value = '') {
|
||||||
|
return JSON.stringify(String(value ?? ''))
|
||||||
|
}
|
||||||
|
|
||||||
|
function parseMetaLine(line = '') {
|
||||||
|
const idx = line.indexOf(':')
|
||||||
|
if (idx < 0) return null
|
||||||
|
const key = line.slice(0, idx).trim()
|
||||||
|
const rawValue = line.slice(idx + 1).trim()
|
||||||
|
if (!key) return null
|
||||||
|
try {
|
||||||
|
return [key, JSON.parse(rawValue)]
|
||||||
|
} catch {
|
||||||
|
return [key, rawValue]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function pickFence(content = '') {
|
||||||
|
const matches = String(content || '').match(/`{3,}/g) || []
|
||||||
|
const maxLen = matches.reduce((max, item) => Math.max(max, item.length), 2)
|
||||||
|
return '`'.repeat(maxLen + 1)
|
||||||
|
}
|
||||||
|
|
||||||
|
export function buildDocBlockValue(attrs = {}) {
|
||||||
|
const docType = normalizeDocType(attrs.docType)
|
||||||
|
const docName = String(attrs.docName || `document.${docType}`)
|
||||||
|
const uploadTime = String(attrs.uploadTime || new Date().toISOString())
|
||||||
|
const collapsed = Boolean(attrs.collapsed)
|
||||||
|
const content = sanitizeDocContent(attrs.content || '')
|
||||||
|
return [
|
||||||
|
`type: ${quoteMeta(docType)}`,
|
||||||
|
`name: ${quoteMeta(docName)}`,
|
||||||
|
`uploadTime: ${quoteMeta(uploadTime)}`,
|
||||||
|
`collapsed: ${collapsed ? 'true' : 'false'}`,
|
||||||
|
'---',
|
||||||
|
content,
|
||||||
|
].join('\n')
|
||||||
|
}
|
||||||
|
|
||||||
|
export function parseDocBlockValue(raw = '') {
|
||||||
|
const normalized = String(raw || '').replace(/\r\n?/g, '\n')
|
||||||
|
const separatorIndex = normalized.indexOf(HEADER_SEPARATOR)
|
||||||
|
const headerText = separatorIndex >= 0 ? normalized.slice(0, separatorIndex) : ''
|
||||||
|
const bodyText = separatorIndex >= 0 ? normalized.slice(separatorIndex + HEADER_SEPARATOR.length) : normalized
|
||||||
|
const attrs = {
|
||||||
|
docType: 'txt',
|
||||||
|
docName: 'document.txt',
|
||||||
|
uploadTime: '',
|
||||||
|
collapsed: false,
|
||||||
|
content: sanitizeDocContent(bodyText),
|
||||||
|
}
|
||||||
|
|
||||||
|
for (const line of headerText.split('\n')) {
|
||||||
|
const parsed = parseMetaLine(line)
|
||||||
|
if (!parsed) continue
|
||||||
|
const [key, value] = parsed
|
||||||
|
if (key === 'type') attrs.docType = normalizeDocType(value)
|
||||||
|
if (key === 'name' && value) attrs.docName = String(value)
|
||||||
|
if (key === 'uploadTime' && value) attrs.uploadTime = String(value)
|
||||||
|
if (key === 'collapsed') attrs.collapsed = value === true || value === 'true'
|
||||||
|
}
|
||||||
|
|
||||||
|
if (!attrs.docName) attrs.docName = `document.${attrs.docType}`
|
||||||
|
return attrs
|
||||||
|
}
|
||||||
|
|
||||||
|
export function buildDocBlockMarkdown(attrs = {}) {
|
||||||
|
const value = buildDocBlockValue(attrs)
|
||||||
|
const fence = pickFence(value)
|
||||||
|
return `${fence}${DOC_BLOCK_FENCE_LANG}\n${value}\n${fence}`
|
||||||
|
}
|
||||||
|
|
||||||
|
export function buildDocContextFence(attrs = {}) {
|
||||||
|
const docType = normalizeDocType(attrs.docType)
|
||||||
|
const content = sanitizeDocContent(attrs.content || '')
|
||||||
|
const fence = pickFence(content)
|
||||||
|
return `${fence}${docType}\n${content}\n${fence}`
|
||||||
|
}
|
||||||
|
|
||||||
|
export function buildLegacyDocBlock(attrs = {}) {
|
||||||
|
const docType = normalizeDocType(attrs.docType)
|
||||||
|
const docName = String(attrs.docName || `document.${docType}`)
|
||||||
|
const uploadTime = String(attrs.uploadTime || new Date().toISOString())
|
||||||
|
const content = sanitizeDocContent(attrs.content || '')
|
||||||
|
return `<doc_type="${docType}" doc_name="${docName}" upload_time="${uploadTime}" collapsed="${Boolean(attrs.collapsed)}">\n${content}\n</doc_end>`
|
||||||
|
}
|
||||||
|
|
||||||
|
export function parseLegacyDocBlock(raw = '') {
|
||||||
|
const match = String(raw || '').match(/^<doc_type="([^"]+)"\s+doc_name="([^"]+)"\s+upload_time="([^"]+)"(?:\s+collapsed="([^"]+)")?>\n?([\s\S]*?)\n?<\/doc_end>$/)
|
||||||
|
if (!match) return null
|
||||||
|
return {
|
||||||
|
docType: normalizeDocType(match[1]),
|
||||||
|
docName: match[2] || 'document.txt',
|
||||||
|
uploadTime: match[3] || '',
|
||||||
|
collapsed: match[4] === 'true',
|
||||||
|
content: sanitizeDocContent(match[5] || ''),
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
export function transformDocBlockMarkdownForClipboard(markdown = '') {
|
||||||
|
const pattern = /(^|\n)(`{3,})llm-file[^\n]*\n([\s\S]*?)\n\2(?=\n|$)/g
|
||||||
|
const replacedFence = String(markdown || '').replace(pattern, (full, prefix, _fence, value) => {
|
||||||
|
const attrs = parseDocBlockValue(value)
|
||||||
|
return `${prefix}${buildDocContextFence(attrs)}`
|
||||||
|
})
|
||||||
|
return replacedFence.replace(/<doc_type="[^"]+"\s+doc_name="[^"]+"\s+upload_time="[^"]+"(?:\s+collapsed="[^"]+")?>[\s\S]*?<\/doc_end>/g, (full) => {
|
||||||
|
const attrs = parseLegacyDocBlock(full)
|
||||||
|
return attrs ? buildDocContextFence(attrs) : full
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
export function stripDocBlockMarkdown(markdown = '') {
|
||||||
|
const pattern = /(^|\n)(`{3,})llm-file[^\n]*\n[\s\S]*?\n\2(?=\n|$)/g
|
||||||
|
return String(markdown || '').replace(pattern, '$1').replace(/\n{3,}/g, '\n\n').trim()
|
||||||
|
}
|
||||||
|
|
||||||
|
export function transformLegacyDocBlocksForExport(markdown = '') {
|
||||||
|
return String(markdown || '').replace(/<doc_type="[^"]+"\s+doc_name="[^"]+"\s+upload_time="[^"]+"(?:\s+collapsed="[^"]+")?>[\s\S]*?<\/doc_end>/g, (full) => {
|
||||||
|
const attrs = parseLegacyDocBlock(full)
|
||||||
|
return attrs ? buildDocBlockMarkdown(attrs) : full
|
||||||
|
})
|
||||||
|
}
|
||||||
|
|
||||||
|
export function transformSpecialDocBlocksToLegacy(markdown = '') {
|
||||||
|
const pattern = /(^|\n)(`{3,})llm-file[^\n]*\n([\s\S]*?)\n\2(?=\n|$)/g
|
||||||
|
return String(markdown || '').replace(pattern, (full, prefix, _fence, value) => {
|
||||||
|
const attrs = parseDocBlockValue(value)
|
||||||
|
return `${prefix}${buildLegacyDocBlock(attrs)}`
|
||||||
|
})
|
||||||
|
}
|
||||||
+20
-11
@@ -36,15 +36,18 @@ export const translations = {
|
|||||||
uploadImg: 'Upload Image',
|
uploadImg: 'Upload Image',
|
||||||
uploadFile: 'Upload File',
|
uploadFile: 'Upload File',
|
||||||
uploadDoc: 'Upload Document',
|
uploadDoc: 'Upload Document',
|
||||||
uploadDocTypeWarning: 'Only txt, docx, pptx, pdf formats are supported.',
|
uploadDocTypeWarning: 'Only txt, json, toml, yaml, docx, pptx, pdf formats are supported.',
|
||||||
uploadDocSizeWarning: 'File size cannot exceed 10MB.',
|
uploadDocSizeWarning: 'File size cannot exceed 10MB.',
|
||||||
uploadDocInBlockWarning: 'Cannot insert document inside an existing document block. Please move cursor outside.',
|
uploadDocInBlockWarning: 'Cannot insert document inside an existing document block. Please move cursor outside.',
|
||||||
uploadDocError: 'Document conversion failed:',
|
uploadDocError: 'Document conversion failed:',
|
||||||
uploadFileTypeWarning: 'Unsupported file type. Supported: doc/docx/ppt/pptx/pdf/zip, images, txt/json.',
|
uploadFileTypeWarning: 'Unsupported file type. Supported: doc/docx/ppt/pptx/pdf/zip, images, txt/json.',
|
||||||
uploadMdTypeWarning: 'Only Markdown (.md) files and image files are supported.',
|
uploadMdTypeWarning: 'Only Markdown (.md) files and image files are supported.',
|
||||||
uploadFileError: 'File upload failed.',
|
uploadFileError: 'File upload failed.',
|
||||||
uploadConvertError: 'File conversion failed.',
|
uploadConvertError: 'File conversion failed.',
|
||||||
enableAI: 'Enable AI',
|
uploadBatchLimit: 'Maximum 10 files at once',
|
||||||
|
uploadSizeLimit: 'File exceeds 50MB limit',
|
||||||
|
uploading: 'Uploading files...',
|
||||||
|
enableAI: 'Enable AI',
|
||||||
disableAI: 'Disable AI',
|
disableAI: 'Disable AI',
|
||||||
insertUrl: 'Insert Image from URL',
|
insertUrl: 'Insert Image from URL',
|
||||||
insert: 'Insert',
|
insert: 'Insert',
|
||||||
@@ -90,15 +93,18 @@ export const translations = {
|
|||||||
uploadImg: '上传图片',
|
uploadImg: '上传图片',
|
||||||
uploadFile: '上传文件',
|
uploadFile: '上传文件',
|
||||||
uploadDoc: '上传文档',
|
uploadDoc: '上传文档',
|
||||||
uploadDocTypeWarning: '仅支持 txt、docx、pptx、pdf 格式的文档',
|
uploadDocTypeWarning: '仅支持 txt、json、toml、yaml、docx、pptx、pdf 格式的文档',
|
||||||
uploadDocSizeWarning: '文件大小不能超过 10MB',
|
uploadDocSizeWarning: '文件大小不能超过 10MB',
|
||||||
uploadDocInBlockWarning: '无法在现有文档块内插入新文档,请将光标移到文档外部',
|
uploadDocInBlockWarning: '无法在现有文档块内插入新文档,请将光标移到文档外部',
|
||||||
uploadDocError: '文档转换失败:',
|
uploadDocError: '文档转换失败:',
|
||||||
uploadFileTypeWarning: '不支持的文件类型。仅支持 doc/docx/ppt/pptx/pdf/zip、图片、txt/json。',
|
uploadFileTypeWarning: '不支持的文件类型。仅支持 doc/docx/ppt/pptx/pdf/zip、图片、txt/json。',
|
||||||
uploadMdTypeWarning: '仅支持 Markdown(.md)和图片文件。',
|
uploadMdTypeWarning: '仅支持 Markdown(.md)和图片文件。',
|
||||||
uploadFileError: '文件上传失败',
|
uploadFileError: '文件上传失败',
|
||||||
uploadConvertError: '文件转换失败',
|
uploadConvertError: '文件转换失败',
|
||||||
enableAI: '启用 AI',
|
uploadBatchLimit: '一次最多上传10个文件',
|
||||||
|
uploadSizeLimit: '文件超过50MB限制',
|
||||||
|
uploading: '正在上传文件...',
|
||||||
|
enableAI: '启用 AI',
|
||||||
disableAI: '禁用 AI',
|
disableAI: '禁用 AI',
|
||||||
insertUrl: '通过 URL 插入图片',
|
insertUrl: '通过 URL 插入图片',
|
||||||
insert: '插入',
|
insert: '插入',
|
||||||
@@ -145,9 +151,12 @@ export const translations = {
|
|||||||
uploadFile: 'Upload File',
|
uploadFile: 'Upload File',
|
||||||
uploadFileTypeWarning: 'Unsupported file type. Supported: doc/docx/ppt/pptx/pdf/zip, images, txt/json.',
|
uploadFileTypeWarning: 'Unsupported file type. Supported: doc/docx/ppt/pptx/pdf/zip, images, txt/json.',
|
||||||
uploadMdTypeWarning: 'Only Markdown (.md) files and image files are supported.',
|
uploadMdTypeWarning: 'Only Markdown (.md) files and image files are supported.',
|
||||||
uploadFileError: 'File upload failed.',
|
uploadFileError: 'File upload failed.',
|
||||||
uploadConvertError: 'File conversion failed.',
|
uploadConvertError: 'File conversion failed.',
|
||||||
enableAI: 'AIを有効化',
|
uploadBatchLimit: 'Maximum 10 files at once',
|
||||||
|
uploadSizeLimit: 'File exceeds 50MB limit',
|
||||||
|
uploading: 'Uploading files...',
|
||||||
|
enableAI: 'AIを有効化',
|
||||||
disableAI: 'AIを無効化',
|
disableAI: 'AIを無効化',
|
||||||
insertUrl: 'URLから画像を挿入',
|
insertUrl: 'URLから画像を挿入',
|
||||||
insert: '挿入',
|
insert: '挿入',
|
||||||
|
|||||||
Reference in New Issue
Block a user