The question brokers keep asking me
In my work advising MLSs and brokerages on data strategy, the question I hear more than any other right now is some version of: how do I keep my listings out of AI?
I have written about the liability brokers carry when their agents use unsanctioned AI tools. Here is the same liability running in the other direction. Your listing data is already inside everyone else’s AI, and you still carry the liability for it.
We created the asset. They built the index.
For more than twenty years, every listing photo, description, price change, open house, and sold record has been crawled, stored, and repurposed by search engines and portals. Brokers and MLSs spend billions collectively creating, curating, verifying, and taking legal responsibility for that data.
The indexers spend nothing to acquire it. They need no IDX (Internet Data Exchange) agreement. They follow no MLS rules. They just crawl.
Here is the asymmetry that should bother every broker. If there is a square footage error, a fair housing problem in the remarks, a misdrawn property line, or a photo that should never have published, the search engine has no liability. The AI has no liability. The listing broker does.
The value of the listing asset moved to the indexer. The liability stayed with the broker.
The MLS was built on a simple idea: we share our listings with each other to sell each other’s listings. It was never built to share our listings with the entire internet so someone else can build a business on them.
Indexing is not hacking. It is reading what we published.
Most brokers picture indexing as someone breaking into the MLS. It is simpler than that.
When your site or your IDX feed shows a listing, your server builds a page that is machine readable on purpose: title tags, Open Graph tags, and schema.org markup like RealEstateListing, Offer, PostalAddress, and GeoCoordinates. We spent a decade making listings machine readable to rank on Google. We made them perfectly structured for AI ingestion at the same time.
Discovery comes from your sitemap.xml, your IDX links, your own site. Parsing reads the metadata. Storage puts the listing in their index, their vector database, their training data. Then a buyer asks an AI what is for sale in 93420 under $800,000, and the answer comes from your data without the buyer ever visiting your site.
This is why opting out of one portal changes nothing. The listing was already indexed from five hundred other places.
The fix is already in the code
Search engines built the opt-out decades ago. It is called noindex, and the compliant crawlers honor it.
- The meta robots tag, <meta name=”robots” content=”noindex, nofollow”>, drops a page from the Google, Bing, and Apple search indexes.
- The X-Robots-Tag HTTP header applies the same control at the server level, which is what MLSs and IDX vendors need for feeds, image directories, and API responses.
- Per-bot directives like <meta name=”googlebot” content=”noindex”> block one crawler and leave the rest alone.
For the AI training crawlers, the documented control is robots.txt. GPTBot, ClaudeBot, Google-Extended, and CCBot all publish robots.txt compliance. A robots.txt block for the training crawlers, plus noindex for the search crawlers, covers both uses of your data: the search index and the training set.
In practice it looks like this. The listing broker’s own site carries index, follow. Every cooperating broker’s IDX or VOW (Virtual Office Website) display of that same listing carries noindex, nofollow, with the AI training crawlers disallowed in robots.txt. The search engine finds the listing once, from the authoritative source: the listing broker. Not five hundred duplicate copies from every IDX site in the state.
Nothing about this breaks IDX or VOW. The listing still displays everywhere it needs to display to sell the home. It just stops being indexed from everywhere.
Will bad actors ignore noindex? Yes, some will. But the companies with the most to lose comply. Their documentation says so.
One listing. One broker. One decision.
Here is the policy change. The listing broker holds the exclusive right to decide if, how, and where their listing is indexed by third party search and AI engines. A cooperating broker may display the listing under IDX or VOW for the purpose of selling it. Displaying it is not the same as authorizing its indexing.
That is not a new idea. It is the listing contract, applied to the internet. The listing broker is the seller’s agent, and deciding how the seller’s home is marketed online is the broker’s fiduciary job. MLSs already write the rules for how listing data may be used, who may display it, and under what terms. This is one more rule, running in the same direction as the broker’s duty to the seller.
Three things change the day this goes live. First, the daily bleeding stops: every new listing is indexed from its source and nowhere else. Second, the chain of custody is restored: the broker who controls the presentation is the broker who carries the liability, which is how it should work. Third, the MLS gets its leverage back: any index that wants a comprehensive, deduplicated set of listings has to license it from the source.
For twenty-five years we optimized our websites for bots. It is time to optimize them for brokers.
One listing. One broker. One decision maker on indexing. That is how the value returns to the people who create the asset and carry the liability for it.