fix: improve Office document processing with Direct track

- Force Office documents (PPTX, DOCX, XLSX) to use Direct track after LibreOffice conversion, since converted PDFs always have extractable text - Fix PDF generator to not exclude text in image regions for Direct track, allowing text to render on top of background images (critical for PPT) - Increase file_type column from VARCHAR(50) to VARCHAR(100) to support long MIME types like PPTX - Remove reference to non-existent total_images metadata attribute This significantly improves processing time for Office documents (from ~170s OCR to ~10s Direct) while preserving text quality. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-30 16:22:04 +08:00
parent 6806fff1d5
commit 87dc97d951
5 changed files with 86 additions and 25 deletions
--- a/backend/app/models/task.py
+++ b/backend/app/models/task.py
@@ -36,7 +36,7 @@ class Task(Base):
    task_id = Column(String(255), unique=True, nullable=False, index=True,
                    comment="Unique task identifier (UUID)")
    filename = Column(String(255), nullable=True, index=True)
-    file_type = Column(String(50), nullable=True)
+    file_type = Column(String(100), nullable=True)
    status = Column(SQLEnum(TaskStatus), default=TaskStatus.PENDING, nullable=False,
                   index=True)
    result_json_path = Column(String(500), nullable=True,