Deep dive
AI + Deterministic Overlay: designing reliable hero images
A design for AI-assisted hero images that keeps typography, technical facts, and production builds under deterministic control.
- AI Infrastructure
- Platform Engineering
- Reliability

A hero image has two jobs. It should make someone want to read the article, and it should look like it belongs to the publication.
An image model can help with the first job. Giving it the entire second job creates a harder problem. Now one generation has to interpret the article, leave space for a title, draw the right typography, preserve technical accuracy, and survive several crops.
That is too much responsibility for one image.
The design I am proposing for my Astro and Keystatic publishing setup separates those responsibilities. AI creates an expressive visual ingredient. Code composes the finished hero. Structured content supplies the facts. A person decides whether the result is worth publishing.
This is an architecture proposal, not a report of a completed implementation. The goal is to make generated artwork useful without making publication identity or build reliability depend on a model’s next response.
The image model should not own the article title
A technical hero is editorial content. A label, a chart, or a convincing dashboard can communicate a claim even when it is only decoration.
That is why this design excludes model-generated titles, architecture labels, code, screenshots, logos, and measurements. An abstract image may suggest layers, flow, or computation. It should not invent the details of the system the article describes.
The ownership boundary is simple. AI owns visual interpretation: metaphor, geometry, texture, and atmosphere. Code owns layout: typography, spacing, framing, safe zones, and crops. Structured data owns titles, topics, dates, and verified labels.
These layers can then change independently. A new artwork candidate does not require a new title treatment. A layout revision does not require rewriting the article. Correcting a component name does not require asking an image model to spell it again.
The finished hero is a composition, not a single generated object.

Each layer has an owner. Regenerating artwork should not rewrite the publication’s identity or its facts.
Generation belongs in authoring, not in the build
From an infrastructure perspective, the most important decision is where generation runs.
I want it in the local authoring workflow. The public site should consume approved files, just as it consumes a photograph or a hand-drawn illustration. Neither the static build nor a visitor’s request should call an image-generation API.
The proposed sequence is:
Article content
→ semantic brief
→ artwork specification
→ generated candidates
→ human selection
→ deterministic composition
→ local asset exports
→ approved content references
→ static buildThis limits the consequences of a provider outage. A quota problem can interrupt creation of a new image, but it should not prevent rebuilding the site with its existing assets. A model update should not silently redraw last month’s cover.
The authoring environment may need credentials, retries, and candidate storage. The published site needs the finished files. Keeping those concerns separate is more valuable to me than making generation happen automatically on every build.

Provider access belongs above the static build boundary. Below it, the site consumes committed content and approved local assets.
Translate meaning before writing the prompt
Sending the whole MDX article straight to an image model mixes two questions: what the image should mean, and how the image should be arranged.
The plan introduces two intermediate objects to keep those questions separate.
VisualBrief describes the subject, concept, metaphor, tone, and things to avoid. For an illustrative article about text-first loading, the brief could look like this:
{
"subject": "text-first loading",
"concept": "content remains stable while secondary layers arrive",
"metaphor": "a stable core surrounded by progressively forming layers",
"tone": "precise, quiet, technical",
"avoid": ["spinner", "browser UI", "fake code"]
}This is a proposed example, not output from an implemented extractor. Its purpose is to make the interpretation visible and editable before generating anything.
ArtworkSpec translates that interpretation into visual constraints:
{
"preset": "systems-editorial",
"layout": "editorial-left",
"visualRegion": "right",
"density": "low",
"background": "light",
"safeZones": ["left"],
"forbidden": ["text", "logos", "UI"]
}The brief can stay the same while the layout changes. The specification can stay the same while the provider changes. Provider-specific prompting still needs an adapter, but that adaptation should not leak into the article schema.
This separation also gives review a useful starting point. If the concept is wrong, fix the brief. If the artwork crowds the title, inspect the composition constraints. Regenerating everything would hide which decision actually failed.
Choose the layout before generating the art
I do not want the publication to rearrange itself around whatever image the model happens to produce.
The plan defines four layout families: editorial left, editorial bottom, minimal, and split. Each gives the artwork a known region and reserves space for the deterministic content.
In the editorial-left example, 52% of the canvas is reserved for title and metadata. The right 48% holds the expressive artwork. Those proportions are design choices, not performance measurements.
The prompt should describe that geometry explicitly. Ask for visual weight on the right, low activity on the left, no text, no interface, and no central focal point that competes with the title.
A prompt is not an enforcement mechanism, though. A model may ignore the boundary or create an object that becomes awkward when cropped. The compositor needs to enforce its text region through clipping or a controlled backdrop, and the reviewer needs to reject artwork that does not fit.
Layout-first generation reduces the ambiguity. It does not remove the need to inspect the result.
SVG is the composition format, not the source of truth
The proposed compositor uses SVG to combine the approved artwork with exact titles, metadata, grids, and verified diagram labels.
That gives code control over title wrapping, font weight, padding, line length, and crop geometry. The AI artwork remains an embedded visual layer. It does not get to move the text or rename the components.
The plan’s initial palette makes these responsibilities visible: paper #F7F5F0, ink #12202D, blue #315EFB for control, and teal #0A8F7A for artwork. Coral is reserved for human decisions, warnings, or exceptions.
The important benefit is reuse. One approved artwork can support a page hero, a 1200 × 630 social card, and a listing image. Each export should recompose the title and metadata for its dimensions rather than blindly crop a wide master and hope the text survives.
Determinism needs more than an SVG file. Reproducible exports require fixed source artwork, composition data, fonts, layout rules, renderer versions, and export settings. The pipeline would need to record those inputs before claiming that a rerun produces the same asset.
SVG controls the composition. The structured content remains the authority for what it says.
Give the editor meaningful choices
The proposed CMS workflow follows the same decision order as the pipeline: analyze the article, edit the brief, choose constraints, explore candidates, compose the exports, and approve the result.
It should not begin and end with a raw prompt box. The useful controls correspond to different kinds of change.
Regenerate artwork keeps the concept and layout but asks for another visual candidate. Reinterpret concept revisits the metaphor while keeping the article. Change composition reuses the approved artwork with different layout rules.
Those actions should not be interchangeable. If the title needs more room, changing the composition may be enough. If the metaphor misrepresents the argument, a new crop cannot fix it.
Automation can prepare choices and render files. Selection remains an editorial decision. Approval of an image should also stay separate from publication of the article; a usable cover does not mean the writing is ready to go live.

Choose the action that fixes the failed decision. A new crop cannot repair a misleading metaphor.
Keep production metadata small and provenance separate
The article should reference its finished images through ordinary frontmatter fields. Generation history belongs in a separate manifest.
That manifest would record the brief version and hash, style preset, layout, provider, prompt version, selected candidate, and compositor version. It should also identify the approved source file and rendering inputs needed to reconstruct the exports.
This keeps the MDX readable while retaining enough history to explain where an image came from.
The approval step should create and validate local files before updating image paths. Ideally, content references, assets, and the manifest would be reviewed and committed together. A partially completed generation must not leave the article pointing to a file that does not exist.
The plan also proposes hashing the VisualBrief to detect possible staleness. If a later article edit changes the brief, the system should flag the hero for review rather than regenerate it automatically.
A hash mismatch is a signal, not an editorial verdict. A corrected typo may leave the meaning unchanged. A revised argument may require new artwork. Stable brief serialization and deliberate brief updates matter if this check is going to be useful.
Alt text needs its own review. The prompt describes the intended image; the final pixels show what actually happened. A candidate description should come from the finished image, then be checked for accuracy and usefulness in the page’s context. A beautiful metaphor is not a reason to add visual claims that are not there.
Prove the composition before building the CMS panel
The full plan includes extraction, house styles, a provider adapter, SVG composition, asset processing, review controls, and Keystatic integration. I would not start by implementing all of that.
The smallest useful version is a CLI workflow with one brief format, one style preset, and one layout. Compose a selected source image, export the assets, review them, and write approved references only after validation.
Then test the system against five real posts or projects. That is a proposed acceptance exercise, not a claim that five images have already been generated. The question is whether they look like one publication while still representing different ideas.
If they do not, adding more presets or CMS buttons will not solve the underlying design problem.
Provider replacement is another useful test. Can a second adapter consume the same brief and specification without changing the MDX contract? It need not produce identical art. It needs to preserve the separation between content meaning, visual constraints, and provider behavior.
Only after those boundaries work would I expose the workflow in the local CMS.
What this approach costs
A controlled pipeline takes more work than asking for one finished image. It introduces schemas, rendering rules, manifests, candidate storage, and an approval step. It also limits visual freedom on purpose.
Those costs are reasonable for a publication that needs consistent identity and technically trustworthy imagery. They may be excessive for a one-off illustration with no text or production workflow around it.
There are practical issues the design still needs to resolve: generation budgets, provider data handling, artwork usage rights, retry behavior, retention of rejected candidates, and validation of image files before embedding them. Provider-neutral objects do not eliminate those concerns. They give them a defined place to live.
The success criteria are modest and testable. The artwork matches the article. Titles remain legible at the intended sizes. Technical labels come from reviewed data. Files exist before references change. Builds consume approved local assets. Humans retain the last editorial decision.
The principle I want to preserve
The image model contributes visual interpretation. It does not decide what the publication is called, what the architecture contains, or whether the site can build today.
Code supplies repeatable composition. Content supplies facts. Humans approve the result.
That is the proposed system: a visual publishing workflow where AI has room to contribute without owning the parts that need to stay exact.