Skip to content

Standalone CDC Scraper - #7

Open
azka2001 wants to merge 1 commit into
MedARC-AI:mainfrom
azka2001:adding-cdc-scraper
Open

Standalone CDC Scraper#7
azka2001 wants to merge 1 commit into
MedARC-AI:mainfrom
azka2001:adding-cdc-scraper

Conversation

@azka2001

@azka2001 azka2001 commented Jul 3, 2026

Copy link
Copy Markdown

Added the file for the CDC scraper based on the Meditron scraper with minor changes to the original file to accommodate changes within the html of the CDC website. Using GROBID via docker as I ran this on windows.

@CLAassistant

CLAassistant commented Jul 3, 2026

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@warner-benjamin

Copy link
Copy Markdown
Collaborator

Thanks for the PR.

This looks fine, but the CDC pages that I spot checked have html versions of the pdf. It would be easier and probably more accurate to process the html to markdown rather than pdf to markdown using lxml like nice or beautifulsoup.

I also don't think we need selenium.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants