Research
Web data in the agent: scrape, crawl, brand, news and page monitors with Context.dev
TL;DR
Context.dev gives the agent the public web as data — any page as markdown or JSON, site maps, small crawls, search, brand profiles and company news — plus monitors that email you when a page changes. Like DataForSEO it runs on platform credentials, so the controls are about cost and about acting on other people's websites.
Connecting it, and what it asks for
- 1Nothing to connect — it runs on the platform's own key, like DataForSEO.
- 2Describe the page or flow you want exercised and what should happen.
- 3Writes (form fills, monitors) stop at the approval gate like any other write.
The authorization asks for nothing from you; usage is metered per run with a per-run ceiling. A first goal worth typing: Walk my signup flow, fill the form with test data, and tell me where it breaks.
What the agent can read
- One page as markdown or HTML, fields extracted to your JSON schema (prices, plans, features), passages relevant to a question, or a product record.
- A site's URLs with titles — the cheap first step before choosing what to read — and a crawl of a section, one page at a time into markdown.
- Ranked web search and a researched answer returned in the JSON shape you ask for, with its sources.
- Brand intelligence — description, logos, colors, socials, address, industry — and a site's design system (colors, typography, button and card CSS) for on-brand creative.
- Company news — launches, funding, partnerships — with source and date.
Eleven read tools plus a gateway to the documented data endpoints. Reads never change anything, so they run without approval — every one is metered.
Two things that act outside the product — and are gated
- Filling a form (
interact_page) — up to five browser actions on a live page, then a check that the result is really there. It is a HIGH-risk step that always waits for approval, is refused on login, checkout and payment pages, and never types passwords, one-time codes, card or identity numbers. The typical use is testing your own lead forms. - Page monitors (
create_monitor) — a standing check that re-reads a page, a sitemap or extracted fields every N hours (at most hourly) and emails the person who asked when it changes. Creating, pausing or deleting one is an approval-gated plan step; an organization out of credits has its monitors paused automatically.
Neither is available to the autonomous loop — the same rule as sending email or deploying a page.
The guardrails
- Metered at what Context.dev reports for each call (its credit count), converted at the platform's rate — never an estimate.
- A per-run budget refuses further calls once a run has spent its cap, and says so.
- Cost knobs clamped: crawl pages (≤25, ≤10 over MCP), search results, map size; researched answers default to the fast mode (10 credits, not 100).
- Private addresses refused before anything is sent, and results trimmed with the dropped amount stated — no silent truncation.
- Monitors belong to your organization: every read, change and delete is checked against an ownership record, and a monitor's webhook is signed and verified before anything it says is trusted.
Real use cases
| Goal you'd type | What the agent does |
|---|---|
| "What does our competitor charge now?" | Maps their site, scrapes the pricing page into a plans table (name, price, currency, period, features) and compares it with yours. |
| "Tell me when their pricing page changes." | Creates a semantic page monitor (after your approval) that ignores banners and dates, and emails you a summary of each real change. |
| "Does the lead form on our demo page still work?" | Reads the form, fills it with test values and submits (after approval), then checks the thank-you state and reports verified true or false. |
| "Make these ads on-brand for acme.com." | Pulls the brand profile and the site's design system, and uses its colors, fonts and tone for the creative. |
| "What has this company announced lately?" | Company news with sources and dates, summarised into what it means for your positioning. |
Frequently asked questions
- Do I need a Context.dev account?
- No. It runs on platform credentials, so there is nothing to connect. Every call is metered to your organization at the credits Context.dev reports.
- Can the agent submit forms on other websites?
- Only as an approval-gated step, only on public forms, never on login, checkout or payment pages, and never with passwords, codes, card or identity numbers. It always verifies the result and says honestly when it cannot.
- How do page monitors notify me?
- Each monitor re-checks its page every N hours (at most hourly; the first run records a baseline) and emails the person who asked when something changes. Ask the agent for your monitor changes at any time, and pause or delete a monitor the same way.
- What does it cost?
- Credits per call: 1 for a page or a site map, 1 per crawled page, 1 per 10 search results or news articles, 10 for a brand profile, styleguide or fast researched answer, and 1 per monitor run (10 for an extract monitor). A per-run budget caps what one task can spend.
- Can an MCP client use it?
- Yes — the family appears on keys whose platform-tools setting includes it. Its writes follow the key's write policy (queued for approval, or pre-authorized), form filling is treated as outbound, and crawls are capped lower over MCP.