Tag
A new framework called TV-Edit combines textual instructions and visual prompts for precise image editing, along with a benchmark TV-Edit-Bench for evaluation. The method achieves better spatial control and semantic faithfulness than existing approaches.