https://www.podbean.com/media/share/pb-ngwwj-1b196bd
Today I’m speaking with Dana Mazia, General Manager of the Bright Initiative by Bright Data, which provides nonprofits, academic institutions, and public bodies with pro bono access to public web data and expertise.
Dana is a lawyer, social entrepreneur, and World Economic Forum Global Shaper with more than 14 years of experience across technology, data, and social impact.
Bright Data is a web-data platform that provides the infrastructure organizations use to collect, process, and analyze publicly available information from across the internet. The Bright Initiative is its public-interest program, making that technology and expertise available to organizations working in academic research, public policy, transparency, and social impact.
Conversations about artificial intelligence typically focus on models: what they can do, what they produce, and whether their outputs are accurate or safe. But behind every model is a less visible infrastructure—the systems through which information is collected, organized, and made usable.
That raises important questions about who can access public web data, what “public” really means, and whether an open internet benefits everyone equally or primarily its most powerful actors. These questions become even more urgent as AI systems move from learning from the web to actively operating within it.
Dana and I discuss the politics of data infrastructure, the ethics of web scraping, the risk of information monopolies, and the legal and ethical questions that arise when information from the public web is collected and operationalized at scale. Ultimately, we ask who should have access to this infrastructure—and who will have the power to shape the AI ecosystem built upon it.