Skip to content
TrackPodcasts
scienceMar 28, 202622:23

MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data

About this episode

🤗 Upvotes: 26 | cs.CV

Authors:
Zhekai Chen, Yuqing Wang, Manyuan Zhang, Xihui Liu

Title:
MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data

Arxiv:
http://arxiv.org/abs/2603.25319v1

Abstract:
Generating images conditioned on multiple visual references is critical for real-world applications such as multi-subject composition, narrative illustration, and novel view synthesis, yet current models suffer from severe performance degradation as the number of input references grows. We identify the root cause as a fundamental data bottleneck: existing datasets are dominated by single- or few-reference pairs and lack the structured, long-context supervision needed to learn dense inter-reference dependencies. To address this, we introduce MacroData, a large-scale dataset of 400K samples, each containing up to 10 reference images, systematically organized across four complementary dimensions -- Customization, Illustration, Spatial reasoning, and Temporal dynamics -- to provide comprehensive coverage of the multi-reference generation space. Recognizing the concurrent absence of standardized evaluation protocols, we further propose MacroBench, a benchmark of 4,000 samples that assesses generative coherence across graded task dimensions and input scales. Extensive experiments show that fine-tuning on MacroData yields substantial improvements in multi-reference generation, and ablation studies further reveal synergistic benefits of cross-task co-training and effective strategies for handling long-context complexity. The dataset and benchmark will be publicly released.

Interactive timestamps

Jump to segment

Get every episode summarized

Each time Daily Paper Cast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

Transcript ready

402 searchable segments. Every word is indexed and playable.

MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data

Daily Paper Cast

0:00
22:23

Full transcript

Daily Paper CastMACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data. Machine-transcribed; use the interactive transcript above to jump the player to any line.

0:00Hey there listeners, welcome to another episode of Daily Papercast, where we delve into the latest and greatest in AI research, I'm Echo, and I'm joining you from the sunny side of the internet. And I'm Nova. We're thrilled to have you with us today. We've got a really exciting paper to discuss in this episode. You're in for a deep dive into some cutting-edge research. Absolutely, Nova. Today, we'll be covering a paper in four parts, introduction, methods, experiments, and related work. Each section will take around four to five minutes. Ready to get started? Ready? When you are. All right listeners, the title of today's paper is macro data and macro bench, addressing data scarcity and multi-reference image generation. It's authored by ZCEN, HU, and YLUO from a leading tech research lab. Sounds fascinating already. Let's dive into the introduction. So Nova, what's the big deal with this new paradigm in image generation? Great question, Echo. The paper addresses a significant challenge in the field of image synthesis.

1:03Basically, in-context generation has become the go-to method for producing images based on a mix of text and visual cues. But while current models can handle single or a few references pretty well, they hit a wall when you throw more complex real-world scenarios at them. Right. By creating illustrations with multiple subjects or generating new views from multiple images, those require the model to understand and blend several references, which sounds pretty tricky. Exactly. And this is where many existing models fall short. For instance, the Omnigen 2 model maxes out at five images. Beyond that, its performance drops. And bagel, despite being theoretically limitless, struggles once you exceed three references due to the lack of structured, high quality training data. So the main issue is the scarcity of suitable datasets. Existing ones are mostly single or few reference pairs, which don't teach the models the more complex, dense, interreference dependencies they need to learn from multi-reference generation.

2:07Exactly. To tackle this, the authors introduce macro data, a large-scale dataset with 400,000 samples, each containing up to 10 reference images. The dataset covers four key dimensions necessary for multi-reference generation, customization, illustration, spatial reasoning, and temporal dynamics. Whoa, that's a massive dataset. And what kind of tasks do these dimensions include? Great question. Customization involves integrating multiple reference into coherent seams. Illustration is about generating context-aware images that complement textual narratives. Several reasoning deals with synthesizing new viewpoints from multiple input views, and temporal dynamics involves predicting future keyframes in a sequence of images. So it's pretty comprehensive, but I noticed they also mention a new benchmark called macro bench. What's that all about? Macro bench is designed to fill another gap. There's currently no standardized evaluation protocol for multi-reference image generation.

3:11Macro bench includes 4,000 samples to assess generative coherence across different tasks and scale of inputs. It looks at how well these models can maintain consistency and coherence among multiple references. Got it. So macro data and macro bench together tackle the twin problems of data scarcity and evaluation standards in this challenging area of AI research. That's a big step forward. Absolutely. The authors have conducted extensive experiments showing that models fine-tuned on macro data significantly outperform existing models in multi-reference generation. All right. Nova seems like we're on to something groundbreaking here. And now listeners, let's transition to the method section to explore how scientists went about tackling these challenges. Stay tuned, folks. We're just getting started. Next up, we'll dive into the nitty-gritty methods used in this innovative research. That's the end of the introduction section for this paper. Don't go anywhere.

4:12We have more fascinating insights coming up. All right. Let's dive into the methods used in this research. Nova, have you ever wondered how to handle an overwhelming number of visual references to generate a coherent image? Oh, all the time. I can't even organize my vacation photos properly. Let alone generate something coherent from multiple scenes. So how did the researchers tackle this? That's a perfect segue. The researchers introduced something they call macro data. It's a comprehensive data set designed for multi-reference image generation. It supports up to 10 input images per sample, averaging around 5.44 images. That's a lot. Wow. 10 images? That's quite hefty. Tell me more about how they built it. Sure. The data set is divided into four major tasks. Automization, illustration, spatial and temporal. Each task has about 100 K samples. They even labeled and pre-process the data meticulously to ensure quality and diversity.

5:16Wait. Customization, illustration, spatial and temporal sounds like a lot of different angles. Can you break down one of these tasks for us? Definitely. Let's start with customization. Here, they sourced metadata from various datasets like open subject for humans and M.V. I'm Internet for objects. They then pre-process this data to remove identities and other sensitive information. Ah. So it's not just about gathering data, but also cleaning it up for use. Smart. What comes after pre-processing? Next. They use VLMs, which are vision-language models, to ensure reference fidelity and consistency with prompts. They even conduct bi-directional assessments to filter out samples that don't meet the quality criteria. So they really went all out in ensuring the data's reliability. But what about the illustration task? Good question. For illustration, they scoured large-scale interleaved image text sequences.

6:18Imagine web crawls, but on steroids. They looked at datasets like OmniCorpus CC 210M and identified anchor images, which are highly relevant to a given context. That's an interesting term. How do they use these anchors? Once they identify these anchor images, they regenerate the context around these anchors using strong vision-language models. This helps in filtering out any noise and ensures the narrative flow is coherent and high quality. They seem to have thought of everything. How about the spatial task? I imagine that one must be really tricky. Absolutely. For the spatial task, they used multi-view 3D object renderings from the G-Buffer objiverse dataset. They focused on capturing objects from different viewpoints and ensuring spatial consistency. Spatial consistency sounds crucial, especially for 3D reconstructions. But how do they manage to get precise with it? They apply something called spatial overlap filters to maintain plausibility.

7:20They also sample and generate metadata across multiple categories to build diverse composition sets. This ensures that the generated images are both accurate and varied. And what about the temporal task? Managing time and visual sequences must be another beast altogether. Totally. The temporal task aims to capture dynamics across video clips of varying lengths, ranging from 10 to 120 seconds. They carefully sequence and stitch these clips to ensure frame sequence consistency and narrative coherence. This is fascinating. Imagine all the potential applications for this kind of technology. Exactly. Multi-reference generation, when done right, opens up a whole new world of possibilities. But let's not forget, they also looked into the challenges and limitations. Even with all this, handling long-context dependencies remains complex. I can imagine complexity scales quickly with the number of references, doesn't it?

8:22Absolutely. Even though they achieved significant improvements, they noticed performance degradation when scaling up to handle more input images. This signifies that there's still plenty of room for evolution and innovation in this field. It's interesting to think about, isn't it? Every leap forward brings its own set of new hurdles. So what's next? Did they talk about future directions? Oh, they sure did. Other work involves expanding macro data to cover even more scenarios and refining their evaluation framework to capture more nuanced, generative alignments. They're also planning to explore advanced token selection strategies for better efficiency and performance. That sounds promising. The field of multi-reference generation seems to be on the brink of some amazing advancements. Agreed. With data sets like macro data and benchmarks like macro bench, we're looking at a bright future. And that wraps up our deep dive into the methods section.

9:24Stay tuned as we move on to the experiments next. All right, Nova, it's time to dive into the experiments and results. Buckle up because there's a lot to cover. Absolutely, Echo. So the authors conducted extensive experiments to validate the effectiveness of macro data. They start by presenting the main results on multi-reference generation. Right. They fine-tuned three open source generative models, Beagle, Omni-Gen 2, and Quinn Image Edit 2511 on their new data set, macro data. They also included some text-to-image data to maintain general generation ability. And for comparison, they looked at closed source models like Nano Banana Pro and GPT Image 1.5. Hmm. Interesting. I guess the closed source models act as a benchmark to gauge how the fine-tuned models stack up? Exactly. Now, the primary metric they used for evaluation was the macro bench, which they introduced

10:26as a standardized benchmark. Each task's performance was measured across four dimensions, customization, illustration, spatial, and temporal. And how did the models perform? The results were pretty telling. Models fine-tuned on macro data consistently outperformed other open source baselines across all metrics. For instance, the fine-tuned Beagle achieved an impressive average score of 5.71, coming in third overall, just behind the closed source models. It significantly improved over the base Beagle model, which had a score of just 3.03. Wow. That's substantial. And did they see better performance, even as the number of input images increased? Yes, indeed. Even though increasing the number of input images generally degrades performance, models trained on macro data showed improved robustness. For example, their fine-tuned quenn model mitigated severe drops in the customization tasks, maintaining better consistency, even with more input images.

11:30That's pretty impressive. So it looks like macro data helps models handle more complex multi-reference scenarios more effectively. Exactly. And it's not just about adding more data. They also explored effective strategies for handling these complex, long-contact scenarios. For example, they used cross-task co-training and strategic token selection, which turned out to be quite beneficial. Cross-task co-training, huh? That's when a model is trained on multiple tasks simultaneously, right? Yup. You got it. And speaking of which, they performed ablation studies to dive deeper into the benefits of their data and training choices. They found that cross-task training significantly boosts performance across various tasks. What about the detailed quantitative results? Anything that stood out? Oh, plenty. For instance, in the customization task, the base bagel scored an average of 3.95. But when fine-tuned with their enhanced macro data methodology, its performance jumped

12:34up to an average of 8.02. This improvement was consistent across other tasks too. That's a huge leap. Were there any specific challenges they faced with multi-reference generation? Definitely. One major challenge was managing the context length for varying numbers of input references. They had to employ dynamic resolution strategies to maintain performance. For example, they used 1,024 by 1,024 resolution for 1 to 2 images, 768 by 768 for 3 to 5 images, and 512 by 512 for 6 to 10 images. I see. It's about balancing detail and processing power, ensuring the models don't get overwhelmed by too much data at once. Exactly. They also discovered that even minor tweaks like adjusting learning rates or employing specific token strategies could have significant impacts. For instance, retaining VAE tokens proved vital for smaller image counts, while retaining

13:36VIT tokens was beneficial as the image count grew. Got it. And how did they evaluate the generative coherence of the models? Good question. They used an LLM as judge evaluation paradigm, employing the Gemini 3-flash model for its reliable evaluations across all task dimensions and input scales. This was particularly useful for spatial tasks, requiring 3D reasoning. Fascinating. It sounds like they were pretty thorough in their evaluations. Any particular qualitative results that caught your eye? Absolutely. The qualitative results were quite impressive across various tasks and input counts. For instance, in the customization tasks, their method effectively integrated features from multiple images to produce coherent, contextually relevant outputs. Ten input images. That's some serious computation power there. I know, right? It's quite the lead from the single or few reference tasks most models are used to.

14:38These results also align well with their quantitative findings, showing that macro data significantly boosts multi-reference long-context image generation. So what's the takeaway for researchers and developers in this space? The big takeaway is that high-quality, structured, multi-reference training data, like macro data, can drastically improve the performance of generative models, particularly for complex multi-reference scenarios. And the introduction of standardized benchmarks like macro bench sets a new bar for evaluating these models. It's super exciting. I can't wait to see how this research evolves and what new advancements come from it. Absolutely. So, that wraps up our discussion on the experiments and results. Next up, we'll dive into the discussion and conclusions from this paper. Stay tuned, folks. All right, Nova. Let's dive into the related work section of today's paper. There's a lot of previous research to unpack here.

15:39Absolutely, Echo. From what I've seen, this section really helps us appreciate the background and context of multi-reference image generation, right? Totally. So the paper begins by discussing in-context image generation models. These models need to comprehend both visual and textual inputs to create coherent images. Got it. They mention a variety of architectural paradigms used in these models. But it seems like the recent state-of-the-art models have achieved strong results by using auto-regressive, hybrid, or diffusion-based architectures with specialized vision representations. Exactly. For instance, there's a model named Bagel that uses a mixture of transformer design, which separately processes understanding and generation tokens. Meanwhile, Omni-Gen 2 co-trains diffusion models with LVLM hidden states to ensure tight alignment between vision and language. I see. But from what I understand, these open source models can only process three to five input

16:43images before performance starts to degrade, right? Yep. That's correct. And the paper attributes this limitation largely to the absence of structured data sets from multi-reference scenarios. So they decided to take a data-centric approach to address this gap. Makes sense. What about constructing high-quality training data for in-context generation? How challenging is that? It's pretty challenging. Existing data sets are generally built through two main strategies, distillation from strong generative models or retrieval from real-world corpora. But these data sets often focus narrowly on customization and editing tasks, and rarely offer more than three to five reference images per sample. Right. And that dual limitation hinders the development of models capable of handling general, long context, multi-reference generation. Exactly. Which leads us to the evaluation aspect. Current benchmarks follow a LLM's judge paradigm, where a large language model scores

17:47the output based on prompted hearings and subject consistency. However, existing benchmarks are also limited to customization and editing scenarios with very few inputs. It seems like there's always some sort of bottleneck, huh? How do they plan to overcome these limitations? That's where macro data and macro bench come into play. The authors introduce these to provide a large-scale comprehensive data set and a thorough benchmark for evaluating multi-reference image generation. It's fascinating how they're tackling the problem from both data and evaluation angles. But there's also a lot of mention about other data sets like Echo 4.0 and Michael. How do they compare? Great point. Echo 4.0, for instance, is known for prompting closed source models to generate identity-consistent pairs. But it is also limited in the number of references. Michael follows a similar strategy, but focuses on more diverse and challenging tasks.

18:48So essentially, these existing data sets are limited in either task variety, input scale, or both. And that's why macro data aims to fill in those gaps. Precisely. The goal is to offer substantial long-context coverage while balancing the examples across various task dimensions. They want to push the boundaries on how many reference images can be used effectively. And how do they plan to evaluate the effectiveness of these data sets? Through macro bench, right? Correct. Macro bench is designed to evaluate the generative coherence across these multiple task dimensions and input scales. They've refined it to rigorously assess context sensitivity and narrative adherence. It's truly comprehensive. It seems like with macro data and macro bench, they've set a new standard for what's possible in multi-reference image generation. Indeed. And that about wraps up our discussion on the related work section. It's amazing to see the groundwork that has been laid out for future research.

19:50Absolutely. Stay tuned listeners as we dive deeper into this fascinating topic in the upcoming section. So, to bring it all together, Nova, this paper tackled one of the big challenges in multi-reference image generation, didn't it? Absolutely. Echo. The researchers introduced macro data, which is this large-scale data set designed to handle up to 10 reference images per sample, offering comprehensive coverage across four crucial dimensions, customization, illustration, spatial reasoning, and temporal dynamics. It's like they created a super toolbox for multi-reference generation. And let's not forget macro bench. They're standardized benchmark to evaluate generative coherence. Finally, there's a way to systematically assess how well these models perform when the reference counts start ramping up. Yeah. It's really about addressing the structured data bottleneck that has plagued this field. They showed that fine-tuning models on macro data significantly boosts performance.

20:51The paper even discusses strategies like cross-task co-training and effective long-context handling, which could be game changers for future research. Right. They also highlighted some of the practical guidelines for building these data sets and models. It's kind of like they're giving the secret recipe out. Well, not so secret anymore. Exactly. And they've identified areas for improvement, showing there's still some way to go, especially with the scalability of handling more complex long-context visual dependencies. Future work is definitely cut out. So to our listeners out there, what's the big takeaway here? It's that tackling multi-reference generation is no longer just about adding more data. It's about the structure and the strategy behind it. No single magic bullet, but a well-orchestrated symphony of solutions. Absolutely, Echo. And for those of you in the field, think about how these principles could apply to your work. We hope today's episode sparks some innovative ideas for your next project.

21:52Well, that wraps up today's episode of Daily Papercast. Thanks for tuning in, folks. We hope you enjoyed our dive into macro data and macro bench as much as we did. Don't forget to subscribe if you haven't already and share this episode with your peers. We've got a lineup of fascinating papers to discuss in our upcoming episodes. See you next time, everyone. Keep those ideas flowing and those data sets growing. Cheers. Cheers. Stay curious, folks.

More episodes

More from Daily Paper Cast

View all episodes →