Create crawlers for each source URL using Python modules (Scrapy, Selenium, Requests, BeautifulSoup, Splash).
Develop and maintain Scrapy pipelines and middlewares to manage crawler output.
Build crawlers for all types of websites, overcoming technical roadblocks.
Manage crawlers to handle IP bans, geolocation bans, CAPTCHAs, and bot‑blocking services.
Write SQL queries for database operations using Python modules such as SQLAlchemy.
Deploy the Python scripts and crawlers to Linux‑based AWS servers.
Skills Must Haves
Strong hands‑on experience in Python programming.
Good experience with scraping libraries such as Requests, BeautifulSoup, Selenium and Scrapy.
Proven 3+ years working with web development frameworks such as Flask, Django, FastAPI, Tornado, and pandas, and building APIs and services using REST.
Experience with any RDBMS and strong SQL knowledge.
Clear understanding of object‑oriented concepts.
Proficient understanding of code versioning tools like Git.
Nice to Have
Familiarity with UI frameworks such as Angular or ReactJS.
Strong unit‑testing and debugging skills.
Knowledge of data engineering platforms such as Airflow.
Experience with Docker.
Knowledge of Agile and Scrum practices; experience with Jira and Confluence is a plus.
Experience with ETL and Celery.
Excellent interpersonal skills and ability to work with a diverse team.
Qualifications
Python
Web Scraping using Requests, BeautifulSoup, Selenium, and Scrapy
Frameworks such as Flask, Django, FastAPI, Tornado, pandas