Bright Data
Search, Crawl and Scrape any site, at scale, without getting blocked
0.5.1Bright Data is a web data platform that provides proxy infrastructure and data collection services at scale. This toolkit enables Arcade agents to search the web, scrape any webpage, and extract structured data from major platforms without getting blocked.
Capabilities
- Web scraping: Fetch any public webpage and receive clean Markdown output, suitable for LLM ingestion or downstream processing.
- Multi-engine search: Query Google, Bing, or Yandex with configurable parameters including result count, country code, and search type (web or images).
- Structured data extraction: Pull typed, schema-consistent data from a curated list of platforms — including Amazon (products, reviews), LinkedIn (people, companies), Instagram (profiles, posts, reels, comments), Facebook (posts, marketplace, reviews), X posts, Zillow listings, Booking.com hotels, YouTube videos, and ZoomInfo company profiles.
Secrets
-
BRIGHTDATA_API_KEY— Your Bright Data API key, used to authenticate all requests. Obtain it from the Bright Data dashboard under Account Settings → API Tokens. You must have an active Bright Data account; a free trial is available but some features require a paid plan. -
BRIGHTDATA_ZONE— The name of the Bright Data proxy zone (also called a "dataset" or "zone") that requests are routed through. Zones are created and managed in the Bright Data control panel under Proxies & Scraping Infrastructure. The zone name must match an existing, active zone in your account (e.g.,residential,serp, or a custom name you assigned). Different zone types affect IP pool, speed, and cost.
See the Arcade secrets guide for how to configure secrets in Arcade, and manage them at https://api.arcade.dev/dashboard/auth/secrets.
Available tools(3)
| Tool name | Description | Secrets | |
|---|---|---|---|
Scrape a webpage and return content in Markdown format using Bright Data.
Examples:
scrape_as_markdown("https://example.com") -> "# Example Page
Content..."
scrape_as_markdown("https://news.ycombinator.com") -> "# Hacker News
..."
| 2 | ||
Search using Google, Bing, or Yandex with advanced parameters using Bright Data.
Examples:
search_engine("climate change") -> "# Search Results
## Climate Change - Wikipedia
..."
search_engine("Python tutorials", engine="bing", num_results=5) -> "# Bing Results
..."
search_engine("cats", search_type="images", country_code="us") -> "# Image Results
..."
| 2 | ||
Extract structured data from various websites like LinkedIn, Amazon, Instagram, etc.
NEVER MADE UP LINKS - IF LINKS ARE NEEDED, EXECUTE search_engine FIRST.
Supported source types:
- amazon_product, amazon_product_reviews
- linkedin_person_profile, linkedin_company_profile
- zoominfo_company_profile
- instagram_profiles, instagram_posts, instagram_reels, instagram_comments
- facebook_posts, facebook_marketplace_listings, facebook_company_reviews
- x_posts
- zillow_properties_listing
- booking_hotel_listings
- youtube_videos
Examples:
web_data_feed("amazon_product", "https://amazon.com/dp/B08N5WRWNW")
-> "{"title": "Product Name", ...}"
web_data_feed("linkedin_person_profile", "https://linkedin.com/in/johndoe")
-> "{"name": "John Doe", ...}"
web_data_feed(
"facebook_company_reviews", "https://facebook.com/company", num_of_reviews=50
) -> "[{"review": "...", ...}]" | 2 |