GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Your Smart TV Has a 200 GB/Month AI Scraping Budget It Never Told You About

The AI training data problem has a residential proxy solution. AI companies need to scrape the live web for pre-training, retrieval, and agent grounding. Cloudflare, DataDome, and HUMAN Security block scraping traffic from known datacenter IP ranges. The workaround is routing requests through residential IP addresses — connections that belong to paying home internet customers and arrive at target sites looking like human traffic.

Bright Data, a data-collection company, markets itself as operating the world’s largest residential proxy network. The supply for that network comes from an SDK embedded in consumer apps, including a significant number of connected TV titles. Security researchers at Include Security have now documented how that SDK works, what it does on user devices, and why smart TVs have become the preferred residential proxy endpoint for AI data harvesting.

Why Smart TVs Beat Phones

Smart TVs are structurally superior residential proxies compared to mobile phones:

  • Always plugged in, no battery to protect
  • Permanently on a high-speed home WiFi connection
  • Operating 24/7 in standby, often unattended
  • No mobile data cap limiting throughput
  • No MDM or mobile EDR software monitoring traffic
  • Consent UI requires navigating legal text via TV remote arrow keys

The Bright Data SDK’s publicly queryable configuration sets max_bw_monthly_wifi: 200,000,000,000 bytes — a 200 GB default monthly WiFi budget per enrolled device. One documented app, Petflix on Roku, discloses in its opt-in dialog that Bright Data will “occasionally use your device’s free resources and IP address to download public web data.” The word “occasionally” does not appear in the 200 GB config.

Partner Scale

Bright Data exposes an unauthenticated partner manifest endpoint. Include Security identified the following from it:

PartnerReach
PlayWorks Digital400+ CTV game titles, ~250M TV homes via Comcast, Sky, Cox, LG, Samsung, Vizio, Roku
CloudTVIntegrated across 125+ TV brands and 15+ OEMs
Viber Media (Rakuten)250M–820M monthly users
Moonfrog Labs (Stillfront)~10M MAU on Teen Patti Gold alone
Longvision Media HK5M OTT users across Hong Kong and Malaysia
SupercentLeading Korean mobile publisher by downloads

PlayWorks alone covers approximately 250 million TV homes through direct integrations with major cable and TV hardware providers. These are not fringe apps. They ship on devices that come pre-loaded in living rooms.

Privacy-policy disclosure is the wrong control surface for a television. Most households will never read it. The in-app consent dialog on a TV requires navigating a legal document with arrow keys on a remote. A typical disclosure says the SDK will “occasionally” use device resources. A 200 GB monthly cap is not “occasional” — it is near-continuous standby consumption.

The FBI issued a formal advisory earlier in 2026 on residential proxy network abuse. Academic measurement going back to 2019 shows these networks are overwhelmingly misused for purposes beyond their stated scope. Most press coverage has focused on illegal supply: botnets like Aisuru and Kimwolf, trojanized apps documented in HUMAN Security’s PROXYLIB disclosure, and pre-infected IoT hardware addressed in Google/Mandiant’s IPIDEA takedown.

The legal supply has received far less scrutiny. Bright Data occupies that gap.

The AI Angle

The connection to AI is not speculative. Bright Data’s own blog names web scraping for machine learning as a primary use case. Krebs on Security reported in October 2025 that a “glut of proxies from Aisuru and other sources is fueling large-scale data harvesting efforts tied to various AI projects.”

Frontier AI companies need continuous access to live web content — current prices, product listings, news articles, public records — for grounding and RAG pipelines. Scraping from datacenter IPs triggers bot detection. Scraping through 250 million residential TV connections does not.

The SDK is legal. The consent is thin. The infrastructure is already in place at scale. The Bright Data partner manifest is public and unauthenticated. Any researcher, regulator, or journalist can fetch it today.

What Changes

Three things would need to shift to close this gap. First, TV platform operators — Roku, LG, Samsung, Vizio — would need to audit and enforce SDK disclosure requirements for apps distributed through their stores, the same way Apple and Google review app permissions. Second, regulators treating residential proxy enrollment as a data processing activity would require opt-in disclosures that meet the bar for meaningful consent, not checkbox disclosure inside a remote-navigated terms screen. Third, AI companies using residential proxy services for training data acquisition would need to disclose that practice in their training data documentation.

None of those things are currently required in the United States. The EU AI Act’s training data transparency provisions may eventually reach this supply chain. They are not yet in force.