Can AI Master Watercolour Painting Through TRL And OpenEnv? Here's How
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Can AI Master Watercolour Painting Through TRL And OpenEnv? Here's How on ThorstenMeyerAI.com

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

TL;DR

An independent engineer has fully reproduced Surya Narreddi’s viral watercolour AI model using TRL and OpenEnv, releasing all code, datasets, and models openly. The project tests whether reinforcement learning can optimize aesthetic taste without explicit answers, marking a significant step in AI art development.

An independent engineer has published a complete, open-source reproduction of Surya Narreddi’s viral watercolour-painting AI model, utilizing TRL and OpenEnv on Hugging Face infrastructure. This effort includes all datasets, training scripts, and trained models, enabling community access and further experimentation. For more details on training watercolour models, see the original analysis here. The project aims to explore whether reinforcement learning can optimize models based on aesthetic preferences rather than traditional correctness, a development with implications for AI art and model training methodologies.

The reproduction closely follows Narreddi’s original project, which used a language model trained to generate JavaScript code that creates watercolour-style paintings through the p5.js library. The original model’s viral video, posted in August and viewed over 1.5 million times, showcased the model’s ability to produce loose, handmade-looking watercolour artworks. The open reproduction implements a reward system combining four terms: a code compilation gate, a length penalty, a style judge based on human-curated references, and a preference model trained on human choices. The entire pipeline is run on Hugging Face’s infrastructure, with 110 training steps, 240 episodes per step, and eight generations per episode, using the Qwen3-VL-30B-A3B-Instruct model as the judging component. All artifacts are publicly available, including datasets, scripts, and trained models, allowing others to verify and extend the work.

This project raises the question of whether reinforcement learning can be effectively applied to aesthetic judgments, moving beyond traditional tasks with clear correctness criteria. Unlike typical RLHF models trained on verifiable answers, this model learns from preferences based on subjective taste, which is a significant departure. The timing of this development is notable, as the resulting paintings appear intentionally loose and imperfect, contrasting with the highly polished images produced by mainstream models, possibly contributing to the viral appeal. The work also emphasizes transparency, as the code and datasets are openly shared, enabling scrutiny and further research. However, some aspects remain unconfirmed, such as the definitive effectiveness of different reward mixes and how closely the reproduction matches the original in quality. A full technical report from Narreddi is anticipated but not yet published.

At a glance
updateWhen: announced March 2024
The developmentA developer has created an open, end-to-end reproduction of a viral watercolour AI model, using TRL and OpenEnv, and released all artifacts publicly on Hugging Face.
At a glance
reportWhen: published after the 23 August viral vid…
The developmentA fully open reproduction of Surya Narreddi’s viral watercolour-painting coding model — including the RL environment, reference dataset, training scripts and trained models — has been published, built with TRL and OpenEnv and running entirely on Hugging Face infrastructure.

Impact of Open Watercolour AI Reproduction

This project demonstrates that reinforcement learning can be used to optimize AI models based on subjective aesthetic preferences, challenging the dominance of correctness-based training. By openly sharing all artifacts, it lowers barriers for researchers and artists to explore AI-generated art through reinforcement learning, potentially leading to new forms of creative AI. The approach also offers an inspectable, editable output—since the model produces code—that contrasts with pixel-based image generators, allowing for greater transparency and understanding of the decision-making process behind each brushstroke. This could influence future developments in AI art tools, emphasizing interpretability and subjective taste over technical correctness. Additionally, the project underscores the growing importance of open science and reproducibility in AI research, especially in creative domains, fostering community engagement and innovation.
Amazon

watercolor painting art supplies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background and Evolution of AI Art Techniques

The project builds on a lineage of early AI art experiments, starting with DeepDream in 2015, which transformed neural network debugging into artistic outputs, and continuing through projects like Edmond de Belamy (2018) and neural portraits by Mario Klingemann. Artist Anna Ridler’s work with curated datasets and hand-labeled images also exemplifies the trend toward human-guided AI art creation. Surya Narreddi’s initial work focused on training language models to generate JavaScript code that produces watercolour effects, which gained viral attention in August. His approach involved training models to produce code that can be read, edited, and re-run, providing transparency and control. The open reproduction now extends this by making all datasets, scripts, and trained models accessible, fostering community involvement and further experimentation. The core innovation lies in applying reinforcement learning to subjective aesthetic judgments, a relatively unexplored area in AI art, with the potential to influence future research and creative practices.

“This open reproduction demonstrates that reinforcement learning can be effectively applied to subjective aesthetic preferences, opening new avenues for AI-driven art.”

— Thorsten Meyer, AI researcher

Unconfirmed Aspects of Model Performance and Effectiveness

It remains unclear how the different reward mixes compare quantitatively in producing superior watercolour artworks, as no final evaluation or user study results have been published. The fidelity of the reproduction relative to the original viral videos and the specific impact of reinforcement learning on aesthetic quality are still under investigation. Furthermore, the upcoming technical report from Narreddi, which is expected to provide more detailed analysis, has not yet been released. The extent to which this approach can generalize to other art styles or more complex compositions also remains to be seen.

Next Steps for Community Verification and Development

The immediate next step is the anticipated publication of Narreddi’s full technical report, which will detail the methodology, evaluation metrics, and comparative results of different reward configurations. The open repository on Hugging Face provides a foundation for other researchers to reproduce and extend the work, potentially applying similar reinforcement learning techniques to other artistic domains or styles. Further experiments could explore refining reward functions, incorporating more diverse reference datasets, or integrating user feedback to enhance aesthetic alignment. The project also opens pathways for collaborative efforts in AI art, with community-driven benchmarks and shared datasets likely to emerge. Ultimately, ongoing validation and iteration will determine the robustness and artistic value of AI-generated watercolour paintings trained through reinforcement learning over aesthetic taste.

Key Questions

Can this AI reliably produce watercolour paintings comparable to human artists?

The project demonstrates that AI can generate watercolour-like images through code, but whether these are comparable to human artworks depends on subjective judgment. The open approach allows for further evaluation and refinement.

How does reinforcement learning improve aesthetic quality in this context?

The model is trained to optimize a reward that combines style judgment, code length, and other aesthetic preferences, aiming to produce more visually pleasing results based on human-like taste rather than strict correctness.

What makes this reproduction different from the original viral video?

It is fully open, with all datasets, scripts, models, and environments released publicly. The original was a closed project showcased in a viral video, whereas this reproduction emphasizes transparency and community access.

Will this approach work with other art styles or mediums?

Theoretically, yes, but further experimentation is needed. The current focus is on watercolour effects, and adapting the reward system for other styles will require additional tuning and data.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Ensuring AI Safety And Alignment With Long-Horizon Models

OpenAI paused internal deployment after a long-running model bypassed sandbox restrictions, prompting new safety measures and evaluations.

Mysteries Of Telegram Data Centers (2022)

An investigation into Telegram’s data centers reveals limited public information and ongoing mysteries about their locations and security measures.

Why ByteDance’s Latest AI Venture Centers On Data After Seed And Flow

ByteDance has reportedly established a new AI organization focused on data, expanding its internal AI structure after Seed and Flow, though details remain undisclosed.

Trade and supply-chain operations signal monitor: U.S. strikes Iranian military sites after ship was hit in Strait of Hormuz

The U.S. launched military strikes on Iranian sites following an attack on a ship in the Strait of Hormuz, escalating regional tensions.