Self-hosted public-web data collection

Run data collection actors on your own infrastructure.

Travinci Actors collects public web data (business websites, news feeds, trends and public social pages) into normalized datasets, reports exactly what it could and could not collect, archives to S3 and exposes everything through an API for your integrations.

Accounts are created by your administrator. There is no public sign-up.

What it does

Normalized datasets

Every actor writes the same record shapes (accounts, posts, articles, pages, comments) with provenance, so downstream systems do not care which platform a record came from.

Coverage reporting

Each run states what was complete, partial, blocked or simply not exposed by the platform, instead of silently returning less data.

S3 archiving

Collected records and post thumbnails can be archived to your own S3-compatible bucket, with the archive state shown on every record.

Schedules & comparisons

Run actors on a schedule, watch sources, and compare runs to see what changed between collections.

Supported collection

Each collector is an “actor”: give it sources, it produces a dataset and a coverage report. All collectors are experimental and labelled with what has been verified.

Business websites

A summary of a company site (name, description, role mailboxes, phones, addresses and social profiles from the site and its contact/about pages), or a full-site crawl within a page budget. Personal-looking email addresses are excluded.

News & RSS feeds

Items from RSS and Atom feeds: title, link, short summary, byline, date and categories. Feed metadata only; article pages and full text are not fetched.

Google Trends

The searches trending right now in a country, from the public Google Trends “Trending now” feed. No keyword histories.

YouTube channels & videos

Channel identity and the channel’s own uploads from its public pages, enriched from each watch page (publish date, public counts). The official YouTube Data API is used when credentials are configured.

Facebook Pages

Page identity and the Page-authored posts its public HTML exposes, with the reaction, comment and share counts shown there.

Instagram accounts

Account metadata and the latest public posts (about 12) from the anonymous profile. Views and shares are not published publicly and need the official API.

LinkedIn company pages

Company profile and its latest public updates from the guest organisation page. Older posts are not available publicly and need the official LinkedIn APIs.

X accounts

An account’s own recent posts and verification metadata, from what X serves to logged-out visitors.

TikTok

Business account identity and the videos exposed in public TikTok page data, where TikTok serves it from your network.

Public data only

Social collectors read only public pages, the same pages a signed-out visitor sees. They never sign in, never bypass login walls or bot protection, and respect robots policies unless an administrator explicitly overrides them for a source. They expose only what platforms publish publicly: metrics such as Instagram views and shares, or LinkedIn posts older than the latest updates, need the platforms’ official APIs.

Built for integrations

  • REST API at /api/v1 to start runs, poll status, read datasets and export results.
  • Scoped bearer keys restricted to chosen actors, features and limits; integration keys never hold administration rights.
  • OpenAPI document at /api/openapi.json.
  • Webhooks signed with HMAC on run completion, so you do not have to poll.
  • Used by Marketing Brain OS as its data collection backend.
curl -X POST "$BASE/api/v1/runs" \
  -H "Authorization: Bearer $TRAVINCI_KEY" \
  -H "Content-Type: application/json" \
  -d '{"actorId":"web-business",
       "input":{"urls":["https://example.com"]}}'

# → 202 { "run": { "id": "…", "status": "QUEUED" } }

Have an account?

Sign in with the email and password set through your activation link.

Sign In