Web Scraping & Automation Engineer
Position: Web Scraping & Automation Engineer Experience: 1–3 Years Location: Remote Employment Type: Full-Time Pay: ₹10,000 – ₹15,000 per month
About the Role
We are looking for a hands-on Web Scraping & Automation Engineer to design, build, and scale reliable data extraction and process-automation systems. The ideal candidate should have strong expertise in Python-based scrapers, headless browser automation, anti-bot handling, scheduling and orchestration, and building clean, structured data pipelines from unstructured web sources.
Responsibilities
- Design and develop scalable, production-grade web scrapers and crawlers.
- Build browser automation workflows using Selenium, Playwright, or Puppeteer.
- Handle dynamic, JavaScript-heavy websites, infinite scroll, and authenticated sessions.
- Implement anti-bot mitigation: proxy rotation, user-agent/fingerprint management, rate limiting, retries, and backoff.
- Design data extraction, parsing, cleaning, deduplication, and validation pipelines.
- Extract structured data from HTML, JSON, XML, PDFs, Excel, and scanned documents.
- Build and maintain scheduled automation jobs using Cron, Celery, or Airflow.
- Automate repetitive business processes and internal workflows (RPA-style automation).
- Integrate third-party APIs and build REST APIs / microservices in Python.
- Store and manage scraped data in SQL and NoSQL databases.
- Implement monitoring, logging, and alerting so broken scrapers are detected early.
- Optimize crawl speed, resource usage, and infrastructure cost.
- Deploy and maintain automation systems on cloud servers.
- Collaborate with product, engineering, and data teams to deliver reliable data feeds.
- Stay up to date with changes in anti-scraping technology and automation tooling.
Required Skills
- 1–3 years of experience in web scraping, crawling, or automation development.
- Strong Python programming skills.
- Hands-on experience with Scrapy, BeautifulSoup, lxml, Requests, and httpx.
- Strong experience with headless browser automation: Selenium, Playwright, or Puppeteer.
- Experience handling anti-bot systems: proxies, rotating IPs, headers, cookies, sessions, and CAPTCHA-solving services.
- Solid understanding of HTTP, REST, XHR/network inspection, cookies, and authentication flows.
- Proficiency with XPath, CSS selectors, and Regex.
- Experience with asynchronous programming (asyncio, aiohttp) for concurrent crawling.
- Data handling with Pandas and standard formats (CSV, JSON, Excel).
- Experience with databases: MySQL/PostgreSQL and MongoDB.
- Experience building REST APIs with FastAPI or Flask.
- Job scheduling and task queues: Cron, Celery, or Airflow.
- Familiarity with Git and CI/CD workflows.
Preferred Skills