Skip to content
Ketryon
Blog

8 min read

Set spending caps before an API bill spirals

How to review spending caps, alerts and shutdown controls for AI, databases, email and other usage-based services.

By Written with AI assistance from the source

Dark clouds over an open sea with a clear horizon
Photo: Ant Rozetsky on Unsplash

Why spending caps have become more important

Usage-based pricing is common for AI, databases, hosting, email, storage and monitoring services. You pay not only for the subscription, but also for things such as requests, traffic, compute time, stored data or messages sent.

There has long been a risk that a bug or unexpected traffic could drive up usage. What has changed is how easy it has become to create systems that carry out cost-incurring work themselves. An AI agent can write code, start background jobs, call models and connect services without every step being reviewed manually.

In October 2026, Simon Willison described the need for hard budget caps by default. His point is that an email about increased usage is not enough when the process continues to incur costs after the warning has been sent. The service needs to be able to stop or limit the work when the cap is reached.

This does not apply only to extreme traffic spikes. Costs can also grow because of a job stuck in a loop, aggressive automatic retries, a public API key, a misconfigured webhook or an agent that keeps working because it never receives a clear stopping condition.

If you are building a product, this means financial safeguards need to be treated as part of the system architecture. They should not be left until the bill already looks unusual.

The difference between alerts and hard limits

The terms budget, limit and quota are often used for different things. Read what the setting actually does, not just what it is called.

A budget alert sends a notification when usage or estimated cost passes a chosen level. The service will normally continue to operate. The alert only helps if the right person sees it and has time to act.

A hard spending cap stops billable usage or pauses affected resources when the limit is reached. It is the clearest form of financial protection, but it can also make the product unavailable.

A usage quota limits a particular resource, such as model processing, data transfer, storage or emails sent. It can protect the cost of that resource without providing a complete cap for the account.

A rate limit restricts how quickly something can be used. It can reduce the impact of an attack or loop, but does not necessarily limit total usage over an entire billing period.

A soft shutdown blocks some functions while allowing others to continue. For example, a database may stop accepting writes while reads still work. This may be better than a complete shutdown, but you need to understand which costs can still arise.

Also check how the provider measures usage. A cap is not necessarily an exact financial boundary. Reporting may be delayed, work that has already started may be completed, and some charges may fall outside the limit. Subscriptions, add-ons, reserved resources, domains or other fixed elements may continue to be charged even when variable usage is stopped.

The question is therefore not just whether a service has a spending cap. You need to know:

  • What does the limit cover?
  • At what level does it apply – account, organisation, project, service or API key?
  • What happens when the limit is reached?
  • How quickly does the action take effect?
  • What usage and charges fall outside it?
  • What is required to restart the service?

Map where costs can grow

Start with a simple inventory of all external services the product can use. Include integrations created automatically by Lovable, an AI agent or a project template. A service may be part of the architecture even if you never created the account yourself.

For each provider, document what drives usage and who owns the billing account.

AI and agents

AI costs are affected by how often the model is called, how much text is processed, which model is used and whether the agent can use tools. A seemingly simple task can trigger further requests through retries, searches, code execution or delegation to other agents.

Check in particular whether the same API key is used in development and production. If the key ends up in client-side code, visitors can use it outside your application. Add the provider's spending cap on top of your own limits for users, workflows and background jobs.

Databases and storage

A database can incur variable costs through compute, data transfer, backups, logs, functions and growing storage. Image files and documents may also be copied, converted or downloaded repeatedly.

Check which parts are covered by the provider's limit. A cap on excess usage may not cover selected compute resources or add-ons that you have activated yourself.

Email and messages

A faulty loop can send the same message repeatedly. An unprotected form can also be used for spam. Cost is only part of the problem – the sender domain's reputation can suffer, and legitimate messages may start being blocked.

Limit both total usage and how often the same recipient or event can trigger a message. Use idempotency, which means that the same event should not be processed more than once simply because it is delivered again.

Hosting and background jobs

Server functions, scheduled jobs, image processing and data transfer can continue to run without anyone visiting the product. Pay particular attention to jobs that read from a queue and then create new work in the same queue.

A spending cap with the hosting provider is useful, but supplement it with time limits, restricted concurrency and a way to disable individual functions without stopping the entire product.

Build protection in layers

The provider's hard cap should be the outermost limit, not the only safeguard. If the entire service has to be shut down, several earlier safeguards have already failed.

Use a combination of the following measures:

  • A hard spending cap with the provider. Enable automatic pausing or blocking where available.
  • Early alerts. Send notifications before the limit is reached so that you can investigate a change without an urgent interruption.
  • Limits within the product. Restrict cost-incurring functions by user, organisation, job or period.
  • Restricted concurrency. Prevent the system from starting an unlimited amount of work at the same time.
  • Time limits and stopping conditions. A job or agent should know when it has to stop.
  • Controlled retries. Delay further attempts and set a clear limit when an external system continues to return errors.
  • Emergency shutdown. Make it possible to disable functions such as AI generation or email without a new deployment.
  • Separate environments. Development and testing should not share unrestricted keys and resources with production.

Also decide what the product should do when an external service is stopped. An AI request can be replaced with a clear error message and the option to try again later. An email can be queued. An upload can be blocked before the file is sent. Planned degradation provides a better experience than unexpected knock-on errors throughout the system.

Alerts must reach someone who can act. Check recipients, spam filters and permissions. Critical alerts should preferably be sent through a channel that does not depend on the same service the alert concerns. If the email provider has been stopped, it is unwise to rely solely on email from it.

Limit who can raise or remove the spending cap. The account should have multi-factor authentication, and changes to billing settings should be traceable in an activity log. An attacker who gains access to an administrator account should not be able to remove both technical and financial safeguards easily.

Test before the protection is needed

A setting that has never been tested is only an assumption. Carry out the test in a separate environment where an interruption will not affect real users or important data.

  • Check that the spending cap is enabled on the correct account and project.
  • Confirm which resources, charges and add-ons it covers.
  • Check whether the cap stops usage or only sends an alert.
  • Trigger controlled test usage and verify that the correct notifications arrive.
  • Check which error the application receives when the service is restricted.
  • Make sure the user sees a clear message.
  • Verify that data remains available after pausing or shutdown.
  • Document how the service is reactivated and who is allowed to do it.
  • Check whether restarting happens automatically or requires manual steps.
  • Add a recurring review of the settings to the product's operations checklist.

Review the list when you change plan, enable a new integration or allow an agent to modify the infrastructure. Providers can change both their pricing models and what their limits cover. A setting that was sufficient when the project was created may mean something different when new services are added.

The aim is not to avoid all risk. The aim is to decide in advance how the product should fail: in a controlled, visible way and with limited financial consequences. If you want to review spending caps and shutdown rules in your own product, there is someone to ask.


Source: Simon Willison

Want to know what your build would take?

Three questions, then you get a clickable sketch for free.

Tell me where you are
← All articles

More from Ketryon

Just need a page or a video?

Two standalone services for when a whole subscription is more than you need — a landing page or a launch video, made by the same person.