How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("image-text-to-text", model="microsoft/AesCode-8B")
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
            {"type": "text", "text": "What animal is on the candy?"}
        ]
    },
]
pipe(text=messages)
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoProcessor, AutoModelForMultimodalLM

processor = AutoProcessor.from_pretrained("microsoft/AesCode-8B")
model = AutoModelForMultimodalLM.from_pretrained("microsoft/AesCode-8B", device_map="auto")
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
            {"type": "text", "text": "What animal is on the candy?"}
        ]
    },
]
inputs = processor.apply_chat_template(
	messages,
	add_generation_prompt=True,
	tokenize=True,
	return_dict=True,
	return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=256)
print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:]))
Quick Links

AesCode-8B

AesCode generates information-rich visual artifacts such as slides, posters, and dashboards as HTML/CSS. The output remains structured, editable, and verifiable, but the task poses a distinct challenge: code models cannot see how layout, hierarchy, and color come together on the canvas.

Image generators offer the opposite strength. They compose visually compelling pages but often misrender text, numbers, and logical relationships. AesCode uses an image generated from the same prompt as an aesthetic reference while following the prompt for the required content.

Reference input alone does not solve the problem. Off-the-shelf vision-language models may copy hallucinated content or ignore the reference layout. AesCode separates semantic requirements from visual cues through graph-structured supervision and decoupled cross-modal rewards.

AesCode-8B starts from Qwen3-VL-8B-Instruct and is trained with cold-start SFT followed by GDPO across seven reward channels.

Infographics generated as editable HTML and CSS by AesCode

Intended Use

AesCode is designed to generate information-rich visual artifacts as structured HTML/CSS. It is suited for creating editable slides, posters, dashboards, and reports, as well as for research on multimodal code generation and verifiable visual design.

Results

We evaluate on 300 infographic samples. Rule averages the programmatic Text, Bound, and Chart checks. Visual averages the Content, Layout, and Style checklist dimensions. Overall is the mean of Rule and Visual. Scores are percentages averaged over three generations per prompt with no selection.

Ref. indicates whether the model receives an image-generated reference together with the prompt. Bold marks the best score in each column and underline the second best.

Model Ref. Text Bound Chart Rule Content Layout Style Visual Overall
GPT-5.5 No 88.67 86.18 82.85 85.90 79.27 68.09 52.80 66.72 76.31
GPT-5.5 Yes 89.22 86.85 81.35 85.80 82.44 89.93 57.91 76.76 81.28
Claude Opus 4.8 No 91.40 85.03 87.69 88.04 84.98 64.88 46.72 65.53 76.78
Claude Opus 4.8 Yes 92.45 84.43 73.75 83.55 86.42 89.76 55.51 77.23 80.39
Qwen3-VL-8B No 58.32 77.23 43.30 59.61 38.33 23.97 13.22 25.17 42.39
Qwen3-VL-8B Yes 64.07 70.99 42.35 59.14 52.90 51.33 29.92 44.72 51.93
Qwen3-VL-32B No 66.00 80.31 41.36 62.55 49.81 34.85 19.34 34.67 48.61
Qwen3-VL-32B Yes 71.19 79.61 50.77 67.19 62.04 64.59 38.43 55.02 61.10
AesCode-8B Yes 94.06 88.36 87.79 90.07 86.41 87.80 53.21 75.80 82.94
AesCode-32B Yes 95.34 97.27 90.37 94.33 85.76 90.58 55.99 77.44 85.89

AesCode-8B improves Visual over its reference-conditioned Qwen3-VL-8B backbone by 31.1 points and Overall by 31.0. Its Overall score of 82.94 exceeds the reference-conditioned GPT-5.5 and Claude Opus 4.8 results, while Rule reaches 90.07. The visual gains do not come at the expense of verifiable correctness.

Style is the shared ceiling for every system. It awards credit only when a design needs no further visual revision before delivery, and no model clears 60.

Usage

The model follows the Qwen3-VL chat interface and requires transformers>=4.57. Supply a detailed content prompt and, when available, a reference image for layout and style guidance.

from transformers import AutoProcessor, AutoModelForImageTextToText

model_id = "microsoft/AesCode-8B"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
    model_id, dtype="auto", device_map="auto"
)

messages = [{
    "role": "user",
    "content": [
        {"type": "image", "image": "reference.png"},
        {"type": "text", "text": "<your content prompt>"},
    ],
}]

inputs = processor.apply_chat_template(
    messages, add_generation_prompt=True, tokenize=True,
    return_dict=True, return_tensors="pt",
).to(model.device)

out = model.generate(
    **inputs, do_sample=True, temperature=0.8, top_p=0.95, max_new_tokens=12000
)
print(processor.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

For batched generation, serve with vLLM and allow two images per prompt:

vllm serve microsoft/AesCode-8B --limit-mm-per-prompt image=2 --max-model-len 24576

Reported scores use three rollouts per prompt at temperature 0.8, top-p 0.95, up to 12,000 output tokens and a 24,576-token context.

The model emits a complete HTML document. Tables use HTML table structures and charts use ECharts specifications, keeping both directly inspectable and verifiable.

The reference image is optional. Reference-conditioned training internalizes visual planning into the policy. Measured on AesCode-8B, withholding the reference at inference costs only 1.00 Visual point, against 19.55 for the Qwen3-VL-8B-Instruct backbone and 10.04 for GPT-5.5. Quality is still highest when a reference is supplied.

The repository provides the training prompt template and a pipeline that turns a short idea into the prompt and reference pair the model expects.

Training

Cold-start SFT. 3,000 demonstrations at learning rate 1e-5.

Reinforcement learning. GDPO over 7,408 prompts for 400 steps, using verl's FSDP-vLLM hybrid engine with no critic and no separately trained reward model. AdamW at a constant 5e-6, no warmup, weight decay 0.01, 128 prompts per step with 8 rollouts each, KL and entropy coefficients both 0.001, rollout temperature 1.0 and top-p 1.0, prompt and response each capped at 8,192 tokens. AesCode-32B uses the same recipe and a longer 520-step run.

The reward is not a single scalar. Each target is represented as a design graph describing the whole canvas, which makes every property individually attributable. From that graph come two complementary reward families: deterministic verifiers score the properties parsable from the code and its rendering, while a VLM judge scores the non-parsable ones using a sample-specific Visual Graph Rubric tied to the graph's elements and relations. These form seven channels: execution, text, boundary, tablechart, layout, whitespace, and design. Each is normalized within its rollout group before aggregation, so a dominant signal cannot drown out a weaker one.

Candidate HTML is scored by rendering it in a sandboxed Playwright browser with external requests blocked, which exports the DOM, computed styles, bounding boxes, console status and a screenshot.

Training code, the verifier stack and the Visual Graph Rubric builder are released at https://github.com/microsoft/AesCode.

Citation

@inproceedings{aescode,
  title     = {AesCode: Aesthetic Code Generation with Decoupled Cross-Modal Rewards},
  booktitle = {Under review},
  year      = {2027}
}

License

Released under the Apache 2.0 license, following the Qwen3-VL backbone.

Downloads last month
6
Safetensors
Model size
9B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for microsoft/AesCode-8B

Finetuned
(634)
this model
Quantizations
1 model

Spaces using microsoft/AesCode-8B 2