A Streamlit-based app with a FastAPI backend for extracting structured data (text, images, tables) from websites and PDFs. Processed data is stored in AWS S3 and rendered in a markdown-standardized format. APIs are deployed on Google Cloud Run Service
docker aws-s3 pdf-converter python3 scrapy diffbot webscraping web-data-extraction diffbot-api beautifulsoup4 pymupdf pdf-document-processor google-cloud-run streamlit azure-document-intelligence doclin
-
Updated
Jan 31, 2025 - Jupyter Notebook