WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing

Hui Zhang Zongkai Liu Liqiang Niu‡ Juntao Liu Han Li Zhen Cao
Wenchao Chen Chengduo Zhao Fandong Meng†

Weixin AI, Tencent

‡ Project Lead    † Corresponding Author

At-a-glance bar charts comparing final image quality, agentic chain quality, and prompt quality
WeAgent-MMGenEdit consistently improves both the agentic process and final image quality across generation and editing. It outperforms similarly sized policies and approaches trillion-parameter agents, while image-side post-training further strengthens the generation performance.

Agentic Image Generation

Knowledge-intensive image generation examples covering basketball, football, tennis, and golf

Agentic Image Editing

Knowledge-intensive image editing examples covering a financial report, GPU infographic, films, and university rankings

Abstract

Image generation and editing models have advanced rapidly, yet remain unreliable when prompts require external world knowledge. Bounded and long-tail parametric knowledge prevents direct or reason-then-generate approaches from recovering the required facts and visual appearances. Existing agentic methods add retrieval, but remain constrained by insufficient visual verification, overloaded policy models, and weak integration of textual and visual evidence. We present WeAgent-MMGenEdit, a full-stack recipe spanning a multimodal harness, scalable data construction, a comprehensive benchmark, and post-training for both the agent policy and image backend. WeAgent-Harness manages persistent evidence and provides dedicated verification and integration tools that organize retrieved evidence into a dense carrier. Our pipeline yields 23K supervised trajectories and 14.7K RL tasks with three-layer verifiable checklists, while WeBench-MMGenEdit evaluates bilingual knowledge-intensive generation and multi-image editing. Together, these components enable a 30B-total / 3B-active policy to outperform similarly sized policy models and approach the performance of a trillion-parameter agent.

Comparison of closed-book, existing agentic, and WeAgent-MMGenEdit paradigms
Figure 3. Three paradigms for multimodal knowledge-intensive image generation and editing. (A) Closed-book generation relies solely on parametric knowledge and suffers from severe factual hallucinations. (B) Existing agentic systems retrieve external evidence but fail to effectively verify and integrate it because of insufficient visual inspection, an overloaded policy model, or weak entity–attribute–layout binding. (C) WeAgent-MMGenEdit verifies retrieved visual evidence and organizes multimodal evidence into a dense rendered carrier that provides both content and layout guidance for final generation, yielding the most accurate results.

One system, end to end

Reliable knowledge-intensive generation is a full-stack problem. We jointly build the runtime, data, benchmark, agent policy, and image backend.

01

Multimodal harness

Persistent evidence, explicit visual verification, and code-based integration keep long-horizon research grounded.

02

Verifiable data

Bilingual multi-hop tasks are paired with aligned checklists for the agentic chain, model input, and final image.

03

Bilingual benchmark

Three isolated evaluation levels measure the agentic chain, generation input, and final image.

04

Agent post-training

Supervised fine-tuning and checklist-grounded RL teach the policy to retrieve, verify, and compose evidence.

05

Image post-training

Multi-reference SFT and reward-driven diffusion RL improve reference conditioning and faithful rendering.

WeAgent-Harness

A persistent multimodal runtime for retrieval, verification, structured integration, and delivery.

Architecture of WeAgent-Harness, showing textual grounding, visual grounding, structured planning, image generation, and trajectory recording
Retrieve

Web and image search build a broad pool of external evidence.

Verify

Focused visual analysis separates inspected references from raw candidates.

Integrate

Code-rendered carriers bind entities, attributes, facts, and layout in pixels.

Deliver

One interface supports generation and editing with ordered references.

WeDataset-MMGenEdit

Bilingual multimodal knowledge becomes verifiable multi-hop tasks, expert trajectories, and targeted post-training data.

WeDataset-MMGenEdit construction pipeline from concept bank and knowledge graph to tasks, checklists, trajectories, and evaluation
864Kaligned bilingual concepts
1.48Mtyped knowledge-graph edges
23Khigh-quality SFT trajectories
14.7Ktargeted RL tasks

WeBench-MMGenEdit

A human-audited bilingual benchmark that gives generation and editing equal weight.

Composition of the 300-case WeBench-MMGenEdit benchmark across languages, task types, domains, reasoning depth, and reference-image count
150 Gen/150 Edit
task type
150 English/150 Chinese
language
Chain/Prompt/Image
evaluation levels

Agentic Post-Training

The policy first learns the complete harness protocol from expert trajectories, then improves evidence acquisition and integration with checklist-grounded reinforcement learning.

Checklist-grounded agent reinforcement learning pipeline
Checklist-grounded agentic RL. Grouped trajectories are evaluated by isolated process and generation-input judges, then optimized asynchronously with failure-aware normalization and GSPO.
Stage I

Agentic Supervised Fine-Tuning

Next-token supervision over policy-generated reasoning, tool calls, and final responses teaches reliable use of WeAgent-Harness from 23K expert trajectories.

Stage II

Checklist-Grounded Agentic RL

Two isolated judges score the agentic process and final multimodal input. Invalid tool trajectories are excluded before group-relative GSPO updates.

Image Editing Post-Training

The image backend learns to consume heterogeneous ordered references and faithfully render factual content, fine-grained text, identities, and preserved regions.

Multi-reference image editing supervised fine-tuning and reinforcement learning pipeline
Two-stage image editing post-training. Multi-reference flow-matching SFT establishes stable conditioning before five-dimensional Diffusion-NFT reinforcement learning refines the output.
Stage I

Multi-Reference Editing SFT

Independent resolution buckets and target-only flow-matching supervision support one to five references with different aspect ratios and semantic roles.

Stage II

Multi-Objective Editing RL

Diffusion-NFT optimizes instruction accuracy, text accuracy, preservation, aesthetics, and relative quality while remaining anchored to the SFT model.

Quantitative Results

Browse the complete WeBench-MMGenEdit results, agentic chain and prompt scores, and evaluations on KnowGen and Mind-Bench.

Final image · Gen21.2 → 47.4

Qwen-Image/Edit backend, from direct generation to the full system.

Final image · Edit24.7 → 50.0

Knowledge-intensive editing improves with agent and image-side post-training.

Agentic chain · Gen35.5 → 59.2

The post-trained policy retrieves, verifies, and integrates evidence more reliably.

Prompt quality · Edit21.8 → 50.9

Verified visual knowledge survives into the final generation request.

Qualitative Results

Across knowledge-rich infographics and multi-reference edits, the complete system preserves more of the facts, identities, constraints, and layout requested by the user.

Figure 9 qualitative comparison on knowledge-intensive generation and editing

BibTeX

@article{zhang2026weagent,
  title={WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing},
  author={Zhang, Hui and Liu, Zongkai and Niu, Liqiang and Liu, Juntao and Li, Han and Cao, Zhen and Chen, Wenchao and Zhao, Chengduo and Meng, Fandong},
  journal={arXiv preprint arXiv:2609.05171},
  year={2026}
}