The prompt-versus-reference distinction is a useful way to debug image results. I treat the prompt as the composition layer: subject, setting, camera distance, lighting, palette, and the relationship between objects. A reference image is better for identity and structure: the shape of a product, the silhouette of a character, the layout of a room, or the texture that should survive an edit. Mixing both roles into a single paragraph often makes it harder to tell what the model actually followed.
For lettering, short text and a clean finishing pass are still the reliable combination. Generate the scene with enough empty space and the right aspect ratio, then add a title or label in an editor when exact spelling matters. If the generator has a text-render mode, compare it with the same prompt in the ordinary mode instead of assuming an upscale will repair the glyphs. Keep a few fixed tests: a three-word sign, a small product label, and a number sequence. That quickly shows whether a model is drawing readable characters or only letter-like texture.
I also find it useful to test changes one variable at a time. Hold the prompt steady while changing the reference strength, then hold that steady while changing resolution or denoise. Save the outputs side by side and inspect edges, hands, repeated patterns, small faces, and shadows. A model that is slightly less dramatic but preserves geometry and makes targeted edits predictable is usually faster to use than one that produces a spectacular first frame.
A practical workflow is to keep a small test sheet with the prompt, seed or variation label, reference strength, aspect ratio, and the reason an output was accepted or rejected. That turns a vague “this looks off” reaction into something repeatable. For product images, inspect straight edges and lettering; for portraits, inspect eyes, hands, hair boundaries, and skin texture; for scenes, inspect repeated windows, rails, and shadows. Those checks reveal more than a gallery thumbnail. For anyone who wants to compare generation and editing in one browser workflow, Photoreal AI — AI image generation/editing is available here: https://photorealistic-ai.com/zh
The broader lesson is that prompting,, reference control, and finishing are separate skills. Use the generator for exploration, use references to stabilize identity and structure, and use an editor for anything that must be exact. That division of labor usually gives cleaner results and makes it easier to diagnose whether a problem came from the prompt, the model, or the final export.