Home/Digital workers/Collector
Collector

One library for the leads and content scattered across platforms

Leads, competitor posts and media that are publicly visible on 6 platforms, collected on a schedule against the keywords and scope you set, then cleaned and deduped into your contact list or media library. What comes back depends on what each platform makes public at the time, and that limit is written into the workflow.

What it delivers

  • Deduped leadsGrouped by platform and tag, straight into your contact list
  • Competitor libraryWhat rivals are posting, ordered by account and date
  • Stored mediaPublicly visible images and video enter the media library for the Publisher to use
  • Logs and failure reasonsWhat couldn't be collected, and why, item by item
Why it exists

You've probably run into these

I'm going platform by platform for leads and end up with a few dozen a day

Jobs run on a schedule against your keywords and scope, and results land in the library. Your team reviews filtered output instead of paging through feeds.

Everyone exports their own spreadsheet, and half of the merged file is duplicates

Records are cleaned and deduped on the way in using identifier fields. Duplicates across jobs and across platforms merge into one entry, with every source kept.

The only way to see what competitors post is to scroll their feeds myself

Name the competitor accounts and their public posts get collected on a schedule, ordered by account and date, so shifts in cadence and topic show up.

The data we pull piles up in a spreadsheet and nobody follows up on it

Leads enter the contact list with their tags, so the Conversation agent can queue follow-up straight away. No exporting between tools.

How it works

From collection scope to leads you can work

Two steps need you: the collection scope, and what the data is for. Scope sets the compliance boundary; purpose decides which library the data enters and who can see it. AI makes neither call.

01
Confirm scope and purpose

You name the platforms, keywords, competitor accounts, regions and fields, and state what this batch of data is for.

Needs you
02
Run to schedule

Jobs run on schedule inside a cloud browser environment or cloud phone, over a proxy in the matching region, at a rate the platform will tolerate.

Automatic
03
Clean, dedupe, tag

Invalid entries drop out, duplicates across platforms merge into one record with sources kept, and tags are applied automatically by platform, region and keyword.

Automatic
04
Confirm storage and routing

Whether leads go to the contact list or stay in the warehouse, and whether media is cleared for publishing, is yours to review before anything moves.

Needs you
05
Store and hand off

Leads go to the contact list for the Conversation agent to work, media goes to the media library for the Publisher, and the full dataset stays in the warehouse for export.

Automatic

Scope is yours to set. The same keyword is collectable to a different degree on each platform and in each region, and local rules on handling personal data differ too. What to collect, and what to use it for, is your compliance call — not something AI should make for you.

Capabilities

What it can do

Multi-platform collection

6 platforms covered. Collection is limited to what a platform makes publicly visible at the time, and the available fields differ by platform.

Lead collection

Collect public accounts and contact details by keyword, region and topic, recording the source of every entry.

Competitor content

Name an account and its public posts are collected on a schedule, ordered by account and date, showing topic choices and frequency.

Cleaning and dedupe

Invalid entries are dropped before storage, and duplicates across jobs and platforms merge into a single record that keeps all its sources.

Tagging

Tags are applied automatically by platform, region and keyword, and you can add your own. Tags follow the lead into the contact list.

Contact list

Approved leads enter the contact list with their source and tags, ready to assign to the Conversation agent.

Media storage

Publicly visible images and video can be stored in the media library for the Publisher to draw on.

Warehouse and export

Every collected record stays in the warehouse. Filter and export it, or connect through the API and webhooks.

Prerequisites

What you need before you start

The Collector is software. The accounts, environments and network it works through are separate. We spell that out so you don't plan against the wrong cost.

Platform accounts
Some platforms only show content to a logged-in session, so those jobs need your own accounts. We don't supply accounts and we don't register them for you.
Cloud environment or cloud phone
Jobs run in an isolated environment, so they don't touch your other accounts. Content that only exists on mobile needs a cloud phone. Billed on actual usage.
Proxy network
Set the exit route by target region so what you collect is the local version. Use our proxy pool or connect your own.
Scope and purpose
You provide the keywords, regions, competitor accounts and fields, and state what the data is for. That's what keeps collection inside your compliance boundary.
Tagging rulesOptional
Automatic tags work without any setup. If leads will be handed to the Conversation agent and worked by category, define a tag set that matches how your business sorts them.
Works with

Who it works with

FAQ

FAQ

Can it collect any data from any platform?

No, and nobody should promise otherwise. Two limits apply. First, only what a platform currently makes publicly visible: content behind a login, behind a friend connection, or already restricted by the platform is out of scope. Second, local rules on handling personal data — collecting and using personal information requires a lawful basis on your side. Visibility rules also change, so a field that's collectable today may not be tomorrow. When something can't be collected, the log records the reason; we don't pad the results with stale data. Before you sign, a consultant walks through what's actually collectable for your target platforms and regions.

Is the collected data accurate?

The data is what the platform page showed, so accuracy depends on the source. Information on a platform can be out of date or simply false, and collection does not verify it. Cleaning only handles malformed entries and duplicates; it makes no judgement about whether content is true. So work a small batch of leads first, confirm the quality, then widen the scope.

Which platforms are supported?

6 platforms: WhatsApp, Telegram, Reddit, X, YouTube and Facebook. What you can collect, and which fields, varies a lot between them — the options shown during setup are what actually applies.

Can the collection rate get an account restricted?

It can, which is why both rate and concurrency are capped and requests are spread across the day instead of bunched together. Keep collection accounts separate from your publishing and conversation accounts, with their own environments and proxies, so the risk stays in one place. If captchas or access restrictions appear, the job slows down or pauses and logs it.

How are duplicate leads handled?

Records are matched on identifier fields, and duplicates across jobs and platforms merge into one entry with every source retained. The same person collected on two platforms becomes a single record marked with both sources, which makes the outreach channel easier to choose.

Can we export the data or feed it into our own systems?

Yes. Filter the results and export them, or push them to your own CRM or data platform through the API and webhooks. The warehouse keeps the full record set, and exporting doesn't change the status of leads already stored.

Collect a batch of leads and judge the quality

A consultant will tell you what is and isn't collectable for your target platforms and regions first, then quote the resources.