Skip to main content
Mindrift

Senior Python Data Scraping Engineer (Freelance)

RemoteMexico, Italy only
Published
Role
Data Engineering
Experience
Senior
Employment
Part-time10–20h/week
up to $25/hr
Check eligibility

Open to MX, IT only. Set where you work from to check your eligibility.

No BS summary

Senior Python data scraping engineer with 5+ years in data engineering, web scraping, automation, or software development. Must handle complex/dynamic sites, anti-bot mechanisms, data cleaning/validation, AWS or equivalent, Docker, and LLM frameworks. Mexico-based remote freelance role, around 10–20 hours/week, English B2+.

Core skills

PythonWeb ScrapingData Extraction

Required skills

BeautifulSoup/SeleniumJavaScriptAJAXAPIsproxiesCSVJSONGoogle SheetsAWSDockerLangChain/OpenRouter

Optional skills

GitHub

Required languages

English Upper-intermediate (B2) or above required.

What you'll do

  • Own end-to-end data extraction workflows across complex websites, ensuring complete coverage, accuracy, and reliable delivery of structured datasets.
  • Leverage available tools and custom workflows to accelerate data collection, validation, and task execution while meeting defined requirement.
  • Ensure reliable extraction from dynamic and interactive web sources, adapting approaches as needed to handle JavaScript-rendered content and changing site behavior.
  • Enforce data quality standards through validation checks, cross-source consistency controls, adherence to formatting specifications, and systematic verification prior to delivery.
  • Scale scraping operations for large datasets using efficient batching or parallelization, monitor failures, and maintain stability against minor site structure changes.

What they require

  • At least 5+ years of relevant experience in data engineering, web scraping, automation, or software development (required).
  • Preferred: Bachelor’s or Master’s Degree in Engineering, Applied Mathematics, Computer Science, or related technical fields is a plus.
  • Candidates should have a strong technical foundation and practical experience with scripting, automation, and data extraction workflows.
  • We are looking for specialists who can solve non-trivial problems, work confidently with modern development tools and technologies, and systematically collect, structure, and validate data from diverse sources.
  • A methodical, detail-oriented approach and the ability to work independently are essential.
  • Strong experience in Python web scraping, including dynamic content (JS, AJAX, infinite scroll) and APIs via proxies.
  • Proven ability to extract data from complex structures (hierarchies, archived pages, inconsistent HTML).
  • Solid background in data cleaning, normalization, and validation, delivering structured datasets.
  • Demonstrated experience handling anti-bot mechanisms and dynamic site structures at scale.
  • Experience with cloud infrastructure and containerization as part of real workflows.
  • Hands-on experience with LLM frameworks applied to automation tasks.
  • Strong attention to detail and commitment to data accuracy.
  • Self-directed work ethic with ability to troubleshoot independently.
  • Preferred: A link to GitHub is a plus.
  • English proficiency: Upper-intermediate (B2) or above (required).
  • For this project, tasks are estimated to require around 10–20 hours per week during active phases, based on project requirements.
  • This is an estimate, not a guaranteed workload, and applies only while the project is active.
  • Strong experience in Python web scraping (BeautifulSoup, Selenium or similar), including dynamic content (JS, AJAX, infinite scroll) and APIs via proxies
  • Proven ability to extract data from complex structures (hierarchies, archived pages, inconsistent HTML)
  • Solid background in data cleaning, normalization, and validation, delivering structured datasets (CSV, JSON, Google Sheets)
  • Demonstrated experience handling anti-bot mechanisms and dynamic site structures at scale
  • Experience with cloud infrastructure (AWS or equivalent) and containerization (Docker) as part of real workflows
  • Hands-on experience with LLM frameworks (LangChain, OpenRouter, or similar) applied to automation tasks
  • Strong attention to detail and commitment to data accuracy
  • Self-directed work ethic with ability to troubleshoot independently
  • Preferred: A link to GitHub is a plus
  • English proficiency: Upper-intermediate (B2) or above (required)
  • Tasks are estimated to require around 10–20 hours per week during active phases, based on project requirements.

Benefits

  • This part-time remote opportunity is ideal for technical professionals with hands-on experience in web scraping, data extraction and processing.
  • On this project, contributors can earn up to $25 per hour equivalent, depending on their level and pace of contribution.
  • On this project, contributors can earn up to $40 per hour equivalent, depending on their level and pace of contribution.

Toloka AI supports frontier model post-training by building domain-specific reinforcement learning environments, tasks, and evaluation frameworks designed by real practitioners. Mindrift, powered by Toloka, connects top domain experts with cutting-edge AI initiatives.

🇮🇳 IndiaAI

What people say about this company

4.0/ 5

up to $25/hr