A crawler identifying itself as OpenAI’s GPTBot spent four days downloading videos from our music project. When we asked AWS to forgive $749.06 in CloudFront charges, an AI-generated refusal arrived in minutes.
The project is a small independent music studio with songs, lyric videos and a modest audience. We used CloudFront, Amazon’s content delivery service, to get the music to listeners quickly. If a song went viral, we wanted the site to keep working. More traffic would mean a bigger bill, and we were prepared for that. We didn’t expect a crawler to be such a big fan.
It started on Tuesday, September 15th, and continued for a couple more days. Normal demand is low at like 1 GB per day. But on this Tuesday, during working hours for that matter, GPTBot/1.4 decided to turn up the music. 10417 GB of media streamed again and again over nearly 40K requests. This AI Bot (or swarm) was having a wonderful time.
Those requests came from four IP addresses in the same network range. We haven’t checked them against OpenAI’s published crawler list, so the GPTBot attribution rests on the name supplied with the traffic.
The expensive mistake was ours. Each video required a signed URL, a link carrying permission to download it. Our server put fresh signed links into every song page. A crawler reading the page’s code could fetch the videos without touching the player. We’d effectively issued our new fan a backstage pass every time it visited.
By 28 September, we’d changed the player to request a signed link only after someone pressed play. We added an AWS WAF firewall rule to block requests identifying as known AI crawlers on the video paths. We also gave the media hostname its own robots.txt rules, because the rules on our main site didn’t cover that separate hostname.
Sadly, we have to cut your streaming time, GPTBot. You’ll now receive a 403, meaning access denied. Browser playback still works. The firewall enforces the block; robots.txt asks cooperative crawlers to stay away.
Our new AI fan was only part of the story. We were just hit with a big bill, and with all the terrifs and economic challenges in front of us right now, every dollar counts. We reached out to AWS to plead our case. We provided the traffic figures, daily charges, and fixes; offered to send the raw logs; and asked why our CloudFront setup wasn’t eligible for its flat-rate plans. We put “Begging for mercy/relief” as the reason for our request.
The answer arrived immediately, marked “generated using AWS Generative AI capabilities”. AWS’s bot had a favourite too: the Customer Agreement. It reminded us that “account holders are responsible for all activities and applicable charges associated with their account” and declined the credit.
If you serve large files, check your page source for working download links, enforce access rules on those files, and inspect your logs to see who’s fetching them. Set alerts for traffic and spending. Budget alerts help, but they follow billing-data updates. You can’t count on a billing email arriving in time to stop a fast download spree.
As for GPTBot, we’re naming it our fan of the month. No human listener has played one of our videos 12,154 times as you have. Devotion like that earns a backstage pass.