DATA&DATA helps luxury brands understand online market dynamics. We aggregate and analyze large-scale data from across the web to provide actionable insights into the pricing, availability, and visibility of high-end consumer goods.
Our core asset is our data: collected at scale from hundreds of sources, checked, normalized, and turned into analytics for some of the world's most iconic luxury brands. We're a small team, which means data collection work here is concrete and immediately useful.
Job Description
As a Web Scraping & Data Collection Intern, you'll be responsible for feeding and monitoring the data that powers our platform. Concretely, you'll:
Configure and run scrapers on new and existing websites using our in-house scraping framework
Investigate target websites to identify undocumented or public APIs, and figure out the most reliable way to collect their data
Write Python scripts and notebooks to collect, extract, and structure data from a wide variety of sources
Run quality checks on collected data: completeness, consistency, detecting when a source breaks or silently changes
Run and monitor existing collection pipelines, and flag or fix issues when a source stops behaving
Answer ad hoc questions on our database with SQL queries (volumes, coverage, anomalies) for internal and client-facing needs
Preferred Experience
You must be enrolled in a school or university able to provide a convention de stage for the full 6 months.
Must-haves
Working knowledge of Python: requests, playwright (or Selenium), plus the basics (loops, functions, files, JSON, virtual environments)
Basic SQL: you can write a SELECT with a WHERE, a GROUP BY, and a simple JOIN without needing to look everything up
Comfort reading HTML and using browser dev tools (inspecting the DOM, reading the network tab)
Git basics for version control
Autonomy: you're comfortable investigating a problem on your own before asking
Fluency in English (written & spoken); French is a bonus
Nice-to-haves
Prior experience with web scraping, in any context (personal projects count)
Familiarity with anti-bot mechanisms, proxies, or headless browsers
Experience with pandas for quick data checks
What You Get
Real, messy, large-scale data from day one — the kind you can't get from a course project
Genuine technical depth on scraping: reverse-engineering APIs, dealing with sites that don't want to be scraped, keeping collection reliable at scale
Autonomy and ownership over the sources you handle
A flat structure, flexible hours, casual dress, no bureaucracy
Recruitment Process
Initial screening – We review your resume and any additional materials you submit
Phone interview – A call to discuss your background, your motivation, and a few basic technical questions
That's it. No take-home assignment, no multi-round process.
Additional Information
Contract Type: Internship (Between 5 and 7 months)
Location: Paris
Education Level: Bachelor's Degree
Occasional remote authorized
Technologies clés
PythonHTMLSQLGitPandas
Conditions en un coup d’œil
Contrat
Stage
Durée
Non précisée
Début
Non précisé
Télétravail
Non précisé
Source
LinkedIn
Suivi de candidature
Connectez-vous pour suivre vos candidatures et garder des notes privées.