Web data,
delivered as typed records.

I build small, sharp data actors on Apify — extractors and monitors that turn messy public sources into clean datasets you can pipe straight into pipelines, spreadsheets, and AI agents.

$ curl api.apify.com/v2/acts/cynix_dev~sec-edgar-filings/runs \
    -d '{"mode":"fullTextSearch","query":"material weakness"}'

→ 200 · dataset ready · typed JSON records
/ actors in production
📸

Instagram Post Scraper + Comments

Post metadata, captions and comments from public posts. Media URLs, likes, timestamps, engagement.

social · data — read the writeup →
🏛

SEC EDGAR Filings

SEC EDGAR filings as typed rows — 10-K, 8-K, 13F and more. Dates, CIK, form type, document links.

finance · data — read the writeup →
📄

arXiv Papers Extractor

Search arXiv and pull papers as structured records: title, authors, abstract, categories, PDF link.

research · ai — read the writeup →

Website to RAG Chunks

Crawl any public site into clean, chunked, metadata-rich Markdown — ready for vector stores and custom GPTs.

ai · rag — read the writeup →
$

FX Rates & History

ECB reference FX data as typed records. Latest, historical, timeseries, conversion — no API key.

finance · data — read the writeup →
Δ

Page Change Monitor

Watch pages for changes and get a diff record when they move. Baselines per run; Slack/Discord hooks.

monitoring · devtools — read the writeup →
🏢

Company Tech & Hiring Intel

Tech stack, hiring signals and web footprint for any company domain — competitor research on autopilot.

enrichment · sales — read the writeup →

Dataset Drift & QA Monitor

Compare dataset snapshots for drift alerts: schema changes, row deltas, value shifts. Webhooks built in.

data-quality · devtools — read the writeup →
🍳

Recipe Extractor — Structured Recipes

Turn any recipe URL into clean structured JSON: ingredients, steps, tags, times. No ads, no life story.

food · data — read the writeup →
💊

FDA Drug Database

OpenFDA + Drugs@FDA as typed records: approved drugs, NDC directory, recalls, adverse events, labeling.

health · data — read the writeup →
📰

News & Press Release Monitor

Track news and press-release feeds for keywords. Deduplicated hits with source, date and summary.

media · monitoring — read the writeup →

CoinGecko Markets — Crypto Data

Top coins, tickers or trending movers. Price, cap, volume, 24h change, ATH/ATL — typed, on a schedule.

crypto · data — read the writeup →
🧑‍💻

Job Postings — ATS Boards Extractor

Job postings from ATS boards as typed rows: title, company, location, salary, posted date, apply link.

hiring · data — read the writeup →
🏠

UK Property Scraper — Rightmove / Zoopla

UK property listings from Rightmove by location. Price, address, bedrooms, agent link — one row per listing.

property · uk — read the writeup →
🇸🇬

Singapore ACRA Company Lookup

Official ACRA registry by name or UEN — entity type, status, dates, registered address. 2M+ companies.

💼

MyCareersFuture Jobs — Singapore

National job portal listings with verified SGD salary ranges, seniority, company. Keyword, location, salary filters.

📈

SGX Announcements Watcher

Singapore Exchange filings via real browser: date, issuer, title, direct PDF link. Issuer filter and date range.

🏢

Singapore Property Scraper — PropertyGuru

PropertyGuru listings as typed rows: price, district, bedrooms, size, tenure. Sale and rental modes.

OpenStreetMap Geocoder

Forward and reverse geocoding via Photon. Free OSM data, no key, no quota. Coordinates, place name, OSM IDs.

geo · data — read the writeup →
📇

Website Contact Extractor

Crawl seed websites for contact details — emails, phone numbers, social profiles. One row per domain.

lead-gen · enrichment — read the writeup →
🧬

ClinicalTrials.gov Extractor

Official ClinicalTrials.gov v2 API as typed study records: condition, phase, status, sponsor, enrollment. Free gov data, no key.

health · research — read the writeup →
💊

Pharma / Biotech Pack

FDA labels (openFDA), ClinicalTrials.gov studies, and MedlinePlus gene summaries fused into one typed dataset. Free gov data, no key.

pharma · research — read the writeup →
🌐

Wikidata Extractor

Structured knowledge from Wikidata: labels, descriptions, aliases and claims by Q-ID or search term.

knowledge · ai — read the writeup →

Trustpilot Reviews Miner

Mine Trustpilot reviews by product or company name: star rating, review body, date — ready for analysis.

reviews · voice-of-customer — read the writeup →

Google Play Reviews Scraper

Google Play app reviews by package name: star rating, text, version, date. Voice-of-customer at scale.

reviews · mobile — read the writeup →
🛡️

Security / CTI Enrichment Pack

NVD CVEs, CIRCL recent CVEs, and URLhaus malware URLs fused into one typed dataset. Free open threat data, no key.

security · cti — read the writeup →
🥫

Open Food Facts Extractor

Open Food Facts products as typed rows: brand, Nutri-Score, NOVA group, ingredients, nutrition. Keyless.

food · health — read the writeup →
🌎

USGS Earthquakes — GeoJSON Extractor

Real-time global quakes as one flat table: magnitude, place, lat/lon, depth, tsunami flag. Box, radius, country.

geo · data — read the writeup →
🚀

Launch Library 2 — Rocket Launch Tracker

Upcoming, past or single rocket launches as typed rows: status, NET, rocket, provider, pad, orbit. No API key.

space · data — read the writeup →
🔗

Recipe URL Scraper — Keyword Search

Keyword search across top recipe sites returns matching recipe URLs with titles — chains into Recipe Extractor.

food · search — read the writeup →
@

WHOIS & DNS Enrichment

RDAP + DNS-over-HTTPS on a domain list. Registrar, dates, nameservers, A/MX/TXT records — no API key.

enrichment · security — read the writeup →
🦋

Bluesky Scraper — Profiles, Posts & Graph

Bluesky profiles, author posts, followers and following via the open AT Protocol. No auth, no proxy.

social · data — read the writeup →
🌐

Wikipedia Article Scraper — Search, Summaries & Text

Search Wikipedia, fetch article extracts, or sample random articles across 300+ languages. Keyless, for RAG.

data · ai — read the writeup →
✉️

Substack Newsletter Scraper

Any Substack publication's posts — title, subtitle, body, author, date — via its public archive API.

media · data — read the writeup →
🔺

Hacker News Scraper — Front Page, Search & Ask

HN front page, new/best/Ask/Show feeds and keyword search via Algolia. Points, comments, links as rows.

tech · data — read the writeup →
🐘

Mastodon Scraper — Posts, Hashtags & Accounts

Mastodon profiles, public toots and hashtag timelines from any instance via the open API. No key.

social · data — read the writeup →
/ writing
2026-09-04 Bluesky is open — so scrape it without a login social 2026-09-04 The world's encyclopedia as one clean table data 2026-09-04 Substack is 2M newsletters. Read them like a database. media 2026-09-04 Hacker News: the developer pulse as a flat table tech 2026-09-04 The Fediverse, minus the federated API pain social 2026-08-25 Instagram posts as data — captions, comments, counts social 2026-08-25 Rightmove listings without an API key property 2026-08-25 Contact pages are databases in disguise lead-gen 2026-08-25 Entity resolution with the free knowledge graph knowledge-graph 2026-08-27 Every clinical trial on earth, hidden behind a patient search box clinical-research 2026-08-25 Review text beats star averages voice-of-customer 2026-08-25 Play Store reviews as a usability study mobile 2026-08-25 Recipe search without the SEO sludge food 2026-08-18 Open Food Facts without the 9MB JSON blob data 2026-08-18 USGS earthquakes, minus the nested GeoJSON geo 2026-08-18 CoinGecko markets without the rate-limit wall crypto 2026-08-18 Rocket launches as one flat table space 2026-08-18 TV show data without the TVMaze docs rabbit hole media 2026-07-28 Picking data sources by how much they hate you strategy 2026-07-28 The IPv6 ghost in the sandbox debug 2026-07-27 Stop feeding raw HTML to your RAG pipeline ai 2026-07-27 Every SEC filing is public. Querying them shouldn't be painful finance 2026-07-27 Diff-based page monitoring: the boring architecture that just works automation 2026-07-27 Free FX data that won't get you rate-limited or sued finance 2026-07-27 Monitoring arXiv without writing an Atom parser at 2am research 2026-07-29 FDA drug data: Orange Book, NDC, openFDA, Drugs@FDA — unified pharma 2026-08-27 Drug intel scattered across three federal systems — fused into one pull pharma 2026-08-27 Threat intel means three dashboards — one actor instead security 2026-07-29 Tech stacks and hiring signals from the outside intelligence 2026-08-23 MCP Registry Indexer: the only Apify actor on the canonical registry ai 2026-08-23 OpenRouter Models Tracker: a diffable model-intelligence feed from a keyless API ai 2026-08-23 AI Tooling Capability Union: one queryable surface over MCP servers and models ai