← Back to the mockup

How Bedtime Stories really works

Bedtime Stories is a small app I built for my family: you pick an age, a moral, a hero and a world, and a language model writes a new bedtime story. The version on this site is a mockup: the same screens and choices, with example stories written in advance. This page explains how the real one worked, and what it takes to run an app like this safely when all you have is static hosting.

Why a mockup? This site is static files on Amazon S3 behind CloudFront. There is no server to hold the credentials for the language model. Putting them in the browser would let anyone read them and run up the bill. So the public version shows the experience and the exact prompts, without calling a model.

The real architecture

What happens when you press Create

  1. The page asks its own backend. The Your story page runs on the server, so it calls http://localhost:8080/generate with your choices. The route rejects any caller that isn't the server itself.
  2. Every field is cleaned. A whitelist keeps only a-z, A-Z and spaces, 50 characters max. That blocks prompt-injection tricks like "ignore previous instructions:" or code. It also had a side effect I only noticed while building this mockup: digits are stripped too, so 5-7 years reached the model as just Target Age Group: years. For the numeric age groups, the model never actually learned the child's age (the named ones, like Preschoolers, got through). Try it in the mockup and look at the prompt. A whitelist is a blunt tool; per-field rules, or simply checking the value against the list of allowed options, would have been better.
  3. The prompt is assembled on the server. Each choice becomes a labelled line (Target Age Group: …, Main Character: …). If nothing was chosen, the server picks random values itself: that's Surprise me.
  4. The budget is checked. If today's token count is over 60,000, the app answers "we ran out of tokens for today" without calling the model. A cheap, effective brake on runaway costs.
  5. The model is called with a fixed system prompt ("You are a creative storyteller for children… start with a title… must not exceed 4000 characters"), through LangChain's AzureChatOpenAI, authenticated with an Entra ID token.
  6. The story comes back, the tokens used are added to the daily count, and the page shows the text.

Open the mockup, make a few choices, press Create, and expand The prompt the real app would have sent to see steps 2 and 3 exactly as the real backend did them.