Senior Data Collection Engineer (Python / Web Scraping)
- Role
- Data Engineering
- Experience
- Senior
- Employment
- Full-time
Open to BR, MX, AR, UY, CO only. Set where you work from to check your eligibility.
No BS summary
Senior Python engineer for large-scale web scraping and data collection. Must know anti-bot mitigation, browser automation, scraping frameworks, Docker/Linux/Git, SQL/NoSQL, REST APIs, and fluent English. Hiring in Brazil, Mexico, Argentina, Uruguay, or Colombia.
Core skills
Required skills
Optional skills
Required languages
Our client is a fast-growing, remote-first B2B SaaS company where large-scale data collection is at the core of the product. As the platform continues to grow globally, they're looking for an experienced Data Collection Engineer to help build and scale the infrastructure behind one of their core product capabilities.
This is a hands-on engineering role focused on designing resilient scraping infrastructure, overcoming sophisticated anti-bot systems, and collecting high-quality data at scale. You'll own the entire lifecycle of large-scale data collection pipelines, ensuring they remain reliable, scalable, and resilient as the product continues to grow.
Responsibilities
- Infrastructure Strategy & Architecture: Architect, build, and maintain the core infrastructure behind our large-scale asynchronous data collection platform.
- Advanced Resilience Engineering: Design, implement, and continuously improve sophisticated anti-blocking strategies, including browser fingerprinting, proxy rotation, CAPTCHA handling, and other techniques required to maintain reliable data collection.
- Core Development: Design, develop, test, and maintain robust scraping components using Python and modern scraping frameworks such as Playwright, Scrapy, Selenium, Requests, and related tools.
- Operational Excellence: Build monitoring, alerting, and logging systems that help identify issues quickly and continuously improve scraper reliability and data quality.
- Data Pipelines & Integrations: Develop and maintain scalable data ingestion pipelines and integrations with internal and external REST APIs.
- DevOps & Automation: Contribute to infrastructure automation using Docker, CI/CD pipelines, Linux environments, and related DevOps practices.
- Collaboration: Work closely with other engineers to improve our scraping platform, establish engineering standards, and mentor less experienced teammates.
Requirements
- Strong commercial experience building high-volume web scraping and data collection systems using Python.
- Deep practical knowledge of anti-bot techniques, including browser fingerprinting, CAPTCHA solving, proxy management, and blocking mitigation.
- Strong understanding of asynchronous programming, browser automation, HTML parsing, HTTP protocols, and REST APIs.
- Hands-on experience with Playwright, Scrapy, Selenium, or similar scraping frameworks.
- Experience working with Docker, Linux, Git, and modern software development practices.
- Familiarity with SQL and NoSQL databases.
- Strong ownership mindset with the ability to independently drive complex technical projects.
- Fluent English communication skills.
Nice to Have
- Experience with advanced asynchronous frameworks (asyncio, Celery, distributed task queues).
- Experience building monitoring and data quality validation systems.
- Experience mentoring engineers or helping technical teams scale.
- Experience using modern AI-assisted development tools (Claude Code, Cursor, Codex, Windsurf, or similar). We value engineers who use AI as an engineering multiplier—guiding, reviewing and orchestrating AI-generated solutions rather than writing every implementation detail manually.
What We Offer
- High ownership and the opportunity to make a measurable impact on a rapidly growing product.
- Remote-first culture with flexible working arrangements.
- Competitive compensation package.
- Personal and professional development through ongoing learning and coaching.
- A collaborative international engineering team solving technically challenging problems.
- Optional office near Berlin at the Wildau Tech University campus.
What you'll do
- Architect, build, and maintain the core infrastructure behind the large-scale asynchronous data collection platform.
- Design, implement, and continuously improve anti-blocking strategies, including browser fingerprinting, proxy rotation, CAPTCHA handling, and other techniques required to maintain reliable data collection.
- Design, develop, test, and maintain robust scraping components using Python and modern scraping frameworks such as Playwright, Scrapy, Selenium, Requests, and related tools.
- Build monitoring, alerting, and logging systems to identify issues quickly and continuously improve scraper reliability and data quality.
- Develop and maintain scalable data ingestion pipelines and integrations with internal and external REST APIs.
- Contribute to infrastructure automation using Docker, CI/CD pipelines, Linux environments, and related DevOps practices.
- Work with other engineers to improve the scraping platform, establish engineering standards, and mentor less experienced teammates.
What they require
- Strong commercial experience building high-volume web scraping and data collection systems using Python.
- Deep practical knowledge of anti-bot techniques, including browser fingerprinting, CAPTCHA solving, proxy management, and blocking mitigation.
- Strong understanding of asynchronous programming, browser automation, HTML parsing, HTTP protocols, and REST APIs.
- Hands-on experience with Playwright, Scrapy, Selenium, or similar scraping frameworks.
- Experience working with Docker, Linux, Git, and modern software development practices.
- Familiarity with SQL and NoSQL databases.
- Strong ownership mindset with the ability to independently drive complex technical projects.
- Fluent English communication skills.
- Preferred: Experience building monitoring and data quality validation systems.
- Preferred: Experience mentoring engineers or helping technical teams scale.
- Preferred: Ability to use AI as an engineering multiplier by guiding, reviewing, and orchestrating AI-generated solutions rather than writing every implementation detail manually.
Benefits
- High ownership and the opportunity to make a measurable impact on a rapidly growing product.
- Remote-first culture with flexible working arrangements.
- Competitive compensation package.
- Personal and professional development through ongoing learning and coaching.
- Collaborative international engineering team solving technically challenging problems.
- Optional office near Berlin at the Wildau Tech University campus.
OnHires' client is an international digital services company that provides marketing, technology, analytics, and operational support to fast-growing global brands, supporting a rapidly growing iGaming project focused on Latin American markets.