DND VTT Token Prompt Guide

Guide · last updated June 2026

Tokens are the most utilitarian prompt type in DND art generation. They are not meant to be beautiful — they are meant to be readable. A good token lets your players tell which figure on the map is the goblin, which is the paladin, and which is the warlord at 400% zoom on a 1080p screen. This guide covers how to prompt for tokens that actually work on a virtual tabletop.

When to use token prompts

Use token prompts whenever you need a top-down or 3/4 overhead view of a character or creature for placement on a virtual tabletop like Roll20, Foundry VTT, or MapTool. Tokens are the visual shorthand that identifies figures on a battle map — they are small, centered, and designed to be recognizable at 40-80 pixels.

Do not use token prompts when you want to show a character in full detail. A token prompt produces a simple, centered silhouette — it cannot show a character's face, expression, or complex costume layers with any fidelity. If you need a portrait for a character sheet, use the portrait prompt type. If you need a battle map illustration, use the scene prompt type.

The practical test: if the image will be displayed at less than 100 pixels wide on a grid, it is a token. If it will be displayed larger than that, it is a portrait, full-body, or scene.

Composition decisions and trade-offs

The defining constraint of a token is size. At 64x64 pixels (a standard Roll20 grid cell), fine details vanish. The prompt must emphasize features that survive extreme downsampling: silhouette shape, color blocking, and a single distinctive element (a hat, horns, a cape) that makes the figure identifiable from adjacent tokens.

Silhouette rule: If you cannot tell what class a character is from their outline alone — a flowing robe for a wizard, a shield silhouette for a paladin, horns for a tiefling — the token prompt has failed. Add silhouette-defining keywords before any other detail.

View angle is the first and most critical decision. "Top-down view" produces a bird's-eye silhouette that works perfectly on grid maps. "3/4 overhead view" adds a slight angle that reveals the face and front of the costume at the cost of some grid alignment. "Isometric view" produces a 3D-style token that looks great in isolation but aligns poorly with flat 2D maps. Pick top-down for standard VTT use; pick 3/4 if you are building tokens for a specific map style that uses angled art.

Background handling is the second critical decision. "Transparent background" is ideal for VTT tokens because it allows the map to show through. "Simple dark background" works if your VTT platform does not support transparency or if you prefer tokens with a solid backing circle. "No background" is ambiguous — some models interpret this as "white background," others as "no background elements." Specify exactly what you want.

Level of detail is the third axis. More detail in the prompt produces more detail in the output, but detail below the pixel threshold of your VTT is wasted. A token that will be 64px wide does not benefit from "intricate filigree on the breastplate" — that detail becomes noise at display size. Prioritize shape and color over texture and fine detail.

Before and after examples

Before

a dwarf fighter token for roll20, top down, with a shield and axe, fantasy style

After

top-down view token of a dwarf fighter, centered single figure on transparent background, squat broad silhouette, round shield on the left arm painted with a copper hammer emblem, battleaxe held vertically in the right hand, iron helmet with a short nose-guard, reddish-brown beard framing the chin, 1:1 square framing, clean readable outline, flat color blocking, VTT token art, no shadow beneath the figure

The "before" prompt uses the word "token" which is ambiguous — the model may produce a coin-like object instead of a character figure. The "after" prompt specifies the exact view angle, background, and the key silhouette-defining features (shield, axe, helmet shape) that make the dwarf identifiable at small size. "Flat color blocking" tells the model to prioritize bold color areas over fine gradients, which produces tokens that read clearly at 64px.

Before

dragonborn paladin token, white scales, glowing sword, holy symbol

After

3/4 overhead view token of a dragonborn paladin, centered figure on transparent background, white-scaled draconic head with a short snout and brow ridges, heavy plate armor with a golden sunburst holy symbol on the chest, a glowing longsword held pointing forward, sweeping tail curling behind, bright warm lighting from above, bold silhouette, clean edges, VTT token illustration, 1:1 aspect ratio

The "before" prompt is a keyword list that omits the view angle and background entirely. The model has to guess how to present a "dragonborn paladin" — it might produce a full portrait, a side view, or a creature illustration. The "after" prompt locks in the view, background, and silhouette, ensuring the output is usable as a token.

Common failure patterns and corrections

Failure 1: Portrait-style output instead of token. When a prompt does not specify a view angle, the model defaults to a portrait or front-facing stance. This produces a nice-looking character image that is useless on a grid map — it is too tall, too detailed, and oriented incorrectly.
Correction: Always include "top-down view" or "3/4 overhead view" as the first phrase in your token prompt. This is the single most important keyword for this prompt type. Without it, nothing else matters.
Failure 2: Over-detailed tokens. Prompts with "intricate details," "highly detailed armor," and "realistic texture" produce tokens that look great at full resolution but become muddy noise at VTT display size. The detail is wasted and actually hurts readability.
Correction: Replace detail-oriented keywords with shape-oriented ones. "Clean readable outline" and "flat color blocking" produce tokens that scale down well. Add "bold silhouette" to reinforce that shape matters more than texture.
Failure 3: Shadow and perspective artifacts. A shadow beneath the token figure looks natural in isolation but creates visual noise on a map, especially when tokens overlap. Perspective rendering (the token figure casting a shadow to one side) breaks the flat-grid aesthetic of VTT maps.
Correction: Add "no shadow beneath the figure" and "flat orthographic projection" to your prompt. These keywords suppress the model's tendency to add naturalistic lighting effects that clash with map grids.

Model-specific guidance

Midjourney

Midjourney struggles with top-down views because its training data heavily favors portraits and landscapes. Use --ar 1:1 for square tokens. Add --style raw to reduce Midjourney's stylistic gloss, which makes tokens harder to read at small sizes. Reinforce the top-down instruction by saying "bird's-eye overhead view" instead of just "top-down" — Midjourney responds better to evocative descriptions than technical terms.

ChatGPT / DALL-E

ChatGPT image generation or DALL-E can be a practical choice for VTT tokens when instruction-following is more important than painterly atmosphere. Be explicit: "The image must be a top-down overhead view showing a single centered figure on a plain background." Background transparency support varies, so plan to remove a solid background in an image editor when necessary.

Stable Diffusion (SDXL / SD3)

Stable Diffusion can produce excellent tokens with the right tag structure. Lead with: top-down view, centered figure, transparent background, VTT token, 1:1 framing. Use the negative prompt aggressively: side view, portrait, face visible, detailed background, shadow, gradient, photorealistic. These negatives block the model's tendency to produce portrait-style output. ControlNet with a depth map can enforce the overhead angle if your setup supports it.

Reviewed prompt template

[view angle] view token of a [race] [class], centered single figure on [background type], [primary silhouette shape], [color-blocking description for armor/costume], [distinctive element: weapon, shield, horns, cape, tail], [secondary silhouette feature], [head or helmet detail], [lighting: overhead/flat/none], clean readable outline, bold silhouette, flat color blocking, 1:1 square framing, VTT token illustration, no shadow beneath the figure

The template prioritizes view angle, silhouette, and background before any character detail. This order ensures the model processes the structural constraints first. Character-specific details (race, class, weapon) come in the middle. Style and formatting instructions anchor the end.

Frequently asked questions

What size should I generate tokens at?
Generate tokens at 512x512 or 1024x1024 pixels, then downscale to your VTT's grid size (typically 64x64 or 70x70). High-resolution generation gives the model room to place features correctly, and the downscaling process naturally simplifies detail — which is exactly what a token needs. Generating directly at 64x64 produces blurry, unusable output.
Can I use a portrait image as a token?
Technically yes, but it will look wrong. A portrait is oriented vertically and shows a face, while a token is oriented from above and shows a silhouette. A portrait-as-token will appear as a tall rectangle on the grid, misaligned with the square cells, and the face will be unreadable at token display size. Use the token prompt type for proper VTT art.
Why does my token have a white background instead of transparent?
Transparency support varies by model and export workflow. If you requested a transparent background and received white or another solid color, use an image editor or background-removal tool to create the final transparent PNG.
Should tokens for monsters look different from character tokens?
Yes. Monster tokens benefit from emphasizing threat cues — larger silhouette, darker colors, angular shapes. Character tokens benefit from being approachable — rounded shapes, clear role indicators (shield, robe, bow). The token guide template works for both; just adjust the silhouette keywords to match the intent.

Related guides