wang.lingxiao
merge
add6977
|
Raw
History Blame
6.47 kB
# README.md header for Hugging Face Spaces
---
title: Advanced Document to Markdown Converter
emoji: 🚀
colorFrom: blue
colorTo: purple
sdk: gradio
sdk_version: 4.44.0
app_file: app.py
pinned: true
license: mit
python_version: 3.11
hardware: cpu-basic
tags:
- document-processing
- markdown
- pdf-converter
- ai-analysis
- mcp-server-track
- mcp-server
- nlp
- ocr
short_description: Convert any document to Markdown with AI-powered analysis
---
# 🚀 Advanced Document to Markdown Converter
Convert documents to Markdown format with AI-powered analysis and advanced features.
## Features
### 📄 Supported Formats
- **PDF** - With OCR support for image-based PDFs
- **Word Documents** (.docx) - Full formatting preservation
- **PowerPoint** (.pptx) - Slide-by-slide conversion
- **Excel** (.xlsx) - Table extraction and formatting
- **Plain Text** (.txt, .md) - Smart formatting detection
- **Rich Text** (.rtf) - Complete formatting support
- **E-books** (.epub) - Chapter and content extraction
### 🧠 AI-Powered Features
- **Structure Analysis** - Intelligent document organization
- **Topic Extraction** - Automatic keyword and topic identification
- **Entity Recognition** - Named entity detection and classification
- **Content Summarization** - AI-generated document summaries
- **Smart Heading Detection** - Context-aware heading hierarchy
### ⚡ Advanced Capabilities
- **Batch Processing** - Process multiple documents simultaneously
- **OCR Integration** - Extract text from images and scanned documents
- **Custom Templates** - Pre-configured output formats
- **Caching System** - Improved performance for repeated processing
- **Progress Tracking** - Real-time processing status
- **Export Options** - Multiple output formats (MD, HTML, PDF)
### 🔧 Technical Features
- **MCP Server** - Model Context Protocol integration
- **Concurrent Processing** - Multi-threaded document handling
- **Memory Optimization** - Efficient large file processing
- **Error Recovery** - Robust error handling and reporting
## Usage
### Single Document Processing
1. Upload your document
2. Configure processing options
3. Click "Process Document"
4. View results in multiple tabs
### Batch Processing
1. Upload multiple documents
2. Enable combination option if needed
3. Process all documents simultaneously
4. Export results as needed
### MCP Integration
This application can be used as an MCP server with Claude AI:
```json
{
"mcpServers": {
"document_converter": {
"command": "npx",
"args": [
"mcp-remote",
"https://YOUR-SPACE-URL/gradio_api/mcp/sse",
"--transport",
"sse-only"
]
}
}
}
```
## Installation
### Local Development
```bash
git clone https://huggingface.co/spaces/YOUR-USERNAME/advanced-document-converter
cd advanced-document-converter
pip install -r requirements.txt
python app.py
```
### Docker Deployment
```dockerfile
FROM python:3.11-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install -r requirements.txt
# Install system dependencies for OCR
RUN apt-get update && apt-get install -y \
tesseract-ocr \
tesseract-ocr-eng \
&& rm -rf /var/lib/apt/lists/*
COPY . .
EXPOSE 7860
CMD ["python", "app.py"]
```
## API Documentation
### Core Functions
#### `process_document(file_path, options)`
Process a single document and convert to Markdown.
**Parameters:**
- `file_path` (str): Path to the document file
- `options` (dict): Processing configuration
- `enable_ai_analysis` (bool): Enable AI-powered analysis
- `include_frontmatter` (bool): Add YAML frontmatter
- `generate_toc` (bool): Generate table of contents
- `use_cache` (bool): Enable result caching
**Returns:**
- Dictionary with markdown content, structure analysis, and metadata
#### `process_multiple_documents(file_paths, options)`
Process multiple documents concurrently.
**Parameters:**
- `file_paths` (list): List of file paths
- `options` (dict): Processing configuration
- `combine_documents` (bool): Merge into single document
- Additional options from single document processing
**Returns:**
- Dictionary with results for each document and optional combined output
### MCP Functions
#### `extract_document_to_md_process_document`
MCP-compatible function for document processing.
**Parameters:**
- `file_path` (str): HTTP/HTTPS URL to document
- `show_prev` (bool): Return preview only
- `show_struct` (bool): Include structure analysis
## Configuration
### Environment Variables
- `MAX_FILE_SIZE_MB` - Maximum file size limit (default: 50)
- `CACHE_DIR` - Directory for cached results
- `WORKERS` - Number of concurrent workers
- `ENABLE_OCR` - Enable OCR processing by default
### Processing Options
- **AI Analysis**: Uses spaCy NLP models for advanced text analysis
- **OCR**: Tesseract-based optical character recognition
- **Caching**: Redis-compatible caching for improved performance
## Dependencies
### Core Requirements
- `gradio>=4.0.0` - Web interface framework
- `python-docx>=1.1.0` - Word document processing
- `PyMuPDF>=1.23.0` - PDF processing
- `python-pptx>=0.6.21` - PowerPoint processing
- `openpyxl>=3.1.0` - Excel file processing
### AI/ML Requirements
- `spacy>=3.7.0` - Natural language processing
- `pytesseract>=0.3.10` - OCR capabilities
- `transformers>=4.30.0` - Advanced AI models
### Optional Features
- `matplotlib>=3.7.0` - Visualization capabilities
- `pandas>=2.0.0` - Data processing
- `scikit-learn>=1.3.0` - Machine learning features
## Performance
### Benchmarks
- **Small files** (<1MB): ~2-5 seconds
- **Medium files** (1-10MB): ~10-30 seconds
- **Large files** (10-50MB): ~30-120 seconds
- **Batch processing**: Linear scaling with concurrent workers
### Memory Usage
- **Base memory**: ~200MB
- **Per document**: ~50-100MB additional
- **OCR processing**: +200-500MB peak usage
## Contributing
1. Fork the repository
2. Create feature branch: `git checkout -b feature-name`
3. Commit changes: `git commit -am 'Add feature'`
4. Push to branch: `git push origin feature-name`
5. Submit pull request
## License
MIT License - see LICENSE file for details.
## Support
- **Issues**: Report bugs and feature requests on GitHub
- **Documentation**: Full API documentation available
- **Community**: Join discussions in the Community tab
---
*Built with ❤️ using Gradio, spaCy, and various document processing libraries*