fix: OCR track table data format and image cropping

egg/OCR

Table data format fixes (ocr_to_unified_converter.py):
- Fix ElementType string conversion using value-based lookup
- Add content-based HTML table detection (reclassify TEXT to TABLE)
- Use BeautifulSoup for robust HTML table parsing
- Generate TableData with fully populated cells arrays

Image cropping for OCR track (pp_structure_enhanced.py):
- Add _crop_and_save_image method for extracting image regions
- Pass source_image_path to _process_parsing_res_list
- Return relative filename (not full path) for saved_path
- Consistent with Direct Track image saving pattern

Also includes:
- Add beautifulsoup4 to requirements.txt
- Add architecture overview documentation
- Archive fix-ocr-track-table-data-format proposal (22/24 tasks)

Known issues: OCR track images are restored but still have quality issues
that will be addressed in a follow-up proposal.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>

This commit is contained in:

egg

2025-11-26 18:48:15 +08:00

parent a227311b2d

commit 6e050eb540

8 changed files with 585 additions and 30 deletions

1

requirements.txt

View File

@@ -69,3 +69,4 @@ pylint>=3.0.0
 # ===== Utilities =====
 python-magic>=0.4.27  # File type detection
 beautifulsoup4>=4.12.0  # HTML table parsing for OCR track

fix: OCR track table data format and image cropping

1 requirements.txt Unescape Escape View File

1

requirements.txt

View File