Web data,
delivered as typed records.
I build small, sharp data actors on Apify — extractors and monitors that turn messy public sources into clean datasets you can pipe straight into pipelines, spreadsheets, and AI agents.
$ curl api.apify.com/v2/acts/cynix_dev~sec-edgar-filings/runs \ -d '{"mode":"fullTextSearch","query":"material weakness"}' → 200 · dataset ready · typed JSON records
Instagram Post Scraper + Comments
Post metadata, captions and comments from public posts. Media URLs, likes, timestamps, engagement.
SEC EDGAR Filings
SEC EDGAR filings as typed rows — 10-K, 8-K, 13F and more. Dates, CIK, form type, document links.
arXiv Papers Extractor
Search arXiv and pull papers as structured records: title, authors, abstract, categories, PDF link.
Website to RAG Chunks
Crawl any public site into clean, chunked, metadata-rich Markdown — ready for vector stores and custom GPTs.
FX Rates & History
ECB reference FX data as typed records. Latest, historical, timeseries, conversion — no API key.
Page Change Monitor
Watch pages for changes and get a diff record when they move. Baselines per run; Slack/Discord hooks.
Company Tech & Hiring Intel
Tech stack, hiring signals and web footprint for any company domain — competitor research on autopilot.
Dataset Drift & QA Monitor
Compare dataset snapshots for drift alerts: schema changes, row deltas, value shifts. Webhooks built in.
Recipe Extractor — Structured Recipes
Turn any recipe URL into clean structured JSON: ingredients, steps, tags, times. No ads, no life story.
FDA Drug Database
OpenFDA + Drugs@FDA as typed records: approved drugs, NDC directory, recalls, adverse events, labeling.
News & Press Release Monitor
Track news and press-release feeds for keywords. Deduplicated hits with source, date and summary.
CoinGecko Markets — Crypto Data
Top coins, tickers or trending movers. Price, cap, volume, 24h change, ATH/ATL — typed, on a schedule.
Job Postings — ATS Boards Extractor
Job postings from ATS boards as typed rows: title, company, location, salary, posted date, apply link.
UK Property Scraper — Rightmove / Zoopla
UK property listings from Rightmove by location. Price, address, bedrooms, agent link — one row per listing.
Singapore ACRA Company Lookup
Official ACRA registry by name or UEN — entity type, status, dates, registered address. 2M+ companies.
MyCareersFuture Jobs — Singapore
National job portal listings with verified SGD salary ranges, seniority, company. Keyword, location, salary filters.
SGX Announcements Watcher
Singapore Exchange filings via real browser: date, issuer, title, direct PDF link. Issuer filter and date range.
Singapore Property Scraper — PropertyGuru
PropertyGuru listings as typed rows: price, district, bedrooms, size, tenure. Sale and rental modes.
OpenStreetMap Geocoder
Forward and reverse geocoding via Photon. Free OSM data, no key, no quota. Coordinates, place name, OSM IDs.
Website Contact Extractor
Crawl seed websites for contact details — emails, phone numbers, social profiles. One row per domain.
ClinicalTrials.gov Extractor
Official ClinicalTrials.gov v2 API as typed study records: condition, phase, status, sponsor, enrollment. Free gov data, no key.
Pharma / Biotech Pack
FDA labels (openFDA), ClinicalTrials.gov studies, and MedlinePlus gene summaries fused into one typed dataset. Free gov data, no key.
Wikidata Extractor
Structured knowledge from Wikidata: labels, descriptions, aliases and claims by Q-ID or search term.
Trustpilot Reviews Miner
Mine Trustpilot reviews by product or company name: star rating, review body, date — ready for analysis.
Google Play Reviews Scraper
Google Play app reviews by package name: star rating, text, version, date. Voice-of-customer at scale.
Security / CTI Enrichment Pack
NVD CVEs, CIRCL recent CVEs, and URLhaus malware URLs fused into one typed dataset. Free open threat data, no key.
Open Food Facts Extractor
Open Food Facts products as typed rows: brand, Nutri-Score, NOVA group, ingredients, nutrition. Keyless.
USGS Earthquakes — GeoJSON Extractor
Real-time global quakes as one flat table: magnitude, place, lat/lon, depth, tsunami flag. Box, radius, country.
Launch Library 2 — Rocket Launch Tracker
Upcoming, past or single rocket launches as typed rows: status, NET, rocket, provider, pad, orbit. No API key.
Recipe URL Scraper — Keyword Search
Keyword search across top recipe sites returns matching recipe URLs with titles — chains into Recipe Extractor.
WHOIS & DNS Enrichment
RDAP + DNS-over-HTTPS on a domain list. Registrar, dates, nameservers, A/MX/TXT records — no API key.
Bluesky Scraper — Profiles, Posts & Graph
Bluesky profiles, author posts, followers and following via the open AT Protocol. No auth, no proxy.
Wikipedia Article Scraper — Search, Summaries & Text
Search Wikipedia, fetch article extracts, or sample random articles across 300+ languages. Keyless, for RAG.
Substack Newsletter Scraper
Any Substack publication's posts — title, subtitle, body, author, date — via its public archive API.
Hacker News Scraper — Front Page, Search & Ask
HN front page, new/best/Ask/Show feeds and keyword search via Algolia. Points, comments, links as rows.
Mastodon Scraper — Posts, Hashtags & Accounts
Mastodon profiles, public toots and hashtag timelines from any instance via the open API. No key.