Jobgether

Senior Data Engineer (Web Scraping)

Jobgether

PythonRequestshttpxBeautifulSoupScrapyPlaywrightSeleniumAWSPySparkTerraform

About the Role

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Data Engineer (Web Scraping) based in India. This is a senior, hands-on opportunity to own the next generation of reliable web-data acquisition within a modern data platform. You will design, build, and operate production-grade Python scrapers and scalable ingestion pipelines that support trusted research and data products. The role focuses on improving the maturity, reliability, and scalability of web-scraping capabilities across a growing data environment. You will establish reusable patterns for extraction, scheduling, storage, monitoring, validation, and failure handling. From investigating new data sources to deploying and supporting production workloads, you will have substantial ownership over the full engineering lifecycle. The role combines deep Python engineering with cloud infrastructure, data pipelines, observability, and practical problem-solving in a remote-first environment. You will work independently while collaborating with a distributed engineering team through code reviews, documentation, and structured development workflows. This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Data Engineer (Web Scraping) based in India. This is a senior, hands-on opportunity to own the next generation of reliable web-data acquisition within a modern data platform. You will design, build, and operate production-grade Python scrapers and scalable ingestion pipelines that support trusted research and data products. The role focuses on improving the maturity, reliability, and scalability of web-scraping capabilities across a growing data environment. You will establish reusable patterns for extraction, scheduling, storage, monitoring, validation, and failure handling. From investigating new data sources to deploying and supporting production workloads, you will have substantial ownership over the full engineering lifecycle. The role combines deep Python engineering with cloud infrastructure, data pipelines, observability, and practical problem-solving in a remote-first environment. You will work independently while collaborating with a distributed engineering team through code reviews, documentation, and structured development workflows. Accountabilities: Own the development, deployment, and ongoing operation of web-scraping and web-data ingestion pipelines. Design and establish scalable web-scraping frameworks with reusable patterns for extraction, scheduling, storage, monitoring, validation, and failure handling. Build, maintain, and improve reliable production scrapers for both new and existing data sources. Investigate websites and determine the most appropriate acquisition method, including APIs, direct HTTP requests, HTML parsing, browser automation, or third-party tooling. Evaluate build-versus-buy options for scraping infrastructure and external services, considering capabilities, reliability, cost, operational complexity, and risk. Ensure web-data acquisition activities appropriately account for internal policies, website terms, robots.txt, access restrictions, privacy, and intellectual-property considerations, escalating unclear situations when required. Diagnose and resolve scraping challenges related to website changes, dynamic content, authentication, sessions, rate limits, concurrency, and other operational constraints. Integrate scraping workloads into scalable data-platform and lakehouse architectures. Improve scheduling, monitoring, storage, validation, and operational support for scraping workloads. Use AI-assisted engineering tools where appropriate while maintaining a thorough understanding of, and accountability for, the code being delivered. Support production workloads through monitoring, debugging, maintenance, and continuous improvement. Contribute to a remote engineering environment through code reviews, documentation, ticket-based workflows, and knowledge sharing. Requirements Demonstrated professional experience building and operating production web-scraping systems at scale . Proven ability to independently take substantial scraping projects from initial investigation through implementation, deployment, and ongoing production support. Strong production-level Python engineering skills, with experience developing maintainable applications rather than standalone scripts. Hands-on experience with scraping technologies such as Requests/httpx, BeautifulSoup, Scrapy, Playwright, or Selenium . Strong practical understanding of HTTP, HTML, APIs, JavaScript-rendered websites, and browser/network behavior. Experience addressing common scraping challenges including pagination, authentication, sessions, retries, rate limiting, concurrency, and proxies. Strong understanding of data pipelines, data quality, and how collected data should be validated, stored, and consumed by downstream systems. Experience deploying, monitoring, and supporting production workloads in a cloud environment. Strong debugging, analytical, and problem-solving abilities, with the judgment to make effective engineering decisions independently. Comfortable working within a remote engineering team and participating in code reviews, documentation, and ticket-based development workflows. Experience with AWS is desirable. Familiarity with lakehouse or data-lake architectures, particularly Apache Iceberg , is a plus. Experience with PySpark or other distributed data-processing technologies is beneficial. Familiarity with Docker and containerized workloads is advantageous. Experience with Terraform or other infrastructure-as-code tools is a plus. Familiarity with Grafana or comparable observability platforms is desirable. Experience operating high-volume or distributed crawling systems is beneficial. Experience evaluating or operating commercial scraping, proxy, or browser-infrastructure services is a plus. Experience implementing automated scraper testing, canary runs, or source-drift detection is desirable. Exposure to legal, compliance, privacy, or data-governance processes related to web-data acquisition is advantageous. Strong ownership, autonomy, documentation, communication, and collaboration skills. Benefits Fully remote position within a remote-first technology team. Opportunity to take ownership of a critical web-data acquisition capability and influence its architecture and operating standards. Senior, hands-on role with substantial autonomy across investigation, engineering, deployment, and production support. Work on scalable data pipelines and modern lakehouse architectures supporting research and data products. Exposure to cloud infrastructure, distributed processing, observability, browser automation, APIs, and production scraping technologies. Opportunity to establish reusable engineering patterns and improve the reliability and scalability of data ingestion. Collaboration with a distributed engineering team through code reviews, documentation, and structured workflows. Environment that supports independent problem-solving, technical ownership, and continuous improvement. Fully remote setup available across the relevant distributed team environment. How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1

You'll be redirected to Jobgether's application page

Job Details

Salary

Not disclosed

Location

India

Job type

Full-time

Category

Data Engineering

Experience

Senior

Posted

Today

Job Highlights

  • Senior level role
  • 100% Remote — open to candidates in India
  • Full-time position

About Jobgether

This job is hosted by Jobgether. Clicking Apply opens their site.

More jobs from Jobgether on RC9

Remote Work Style

Mixed

Mix of flexible and scheduled meetings

Your Match

See how well your skills line up with this role, and what you're missing.

AI Cover Letter

Generate a cover letter tailored to this job from your profile.