← Back
AI on a budget

AMD mini PC generates 3D models, images, and speech

A mini PC running Lemonade AI Server can now generate 3D models, images, and narration from a single request without using the cloud.
By
Multi-functional mini PC sits on desk with dual monitors, keyboard, and mouse.
Foto: Symbolbild | notebookcheck.com · Symbolbild (thematisch gesucht: This AMD Strix Halo mini PC generates 3D models images and s) - nicht das Originalfoto der Quelle.
The essentials
  • A Ryzen AI Max+ 395 mini PC with the RPG-HaloTales-V1 model can create a scene with narration, imagery, and audio from a single request.
  • The system uses four different models across four backends to handle text, images, speech, and listening.

I've been testing the Ryzen AI Max+ 395 mini PC for several months. Until recently, I would have described it as nothing more than a machine for serving AI models. It became my go-to local server for AI operations when support for Nvidia was added. From that moment, it operated with a reliable, unchanging efficiency. That changed a few weeks ago when version 11 was released. It introduced major new features that transformed the server beyond a simple chat alternative. Suddenly, it was capable of text-to-speech. It had an image-to-3D conversion system. It had intelligent model routing.

I was curious about this update. I explored the new capabilities. The results exceeded expectations. During a single request to describe a tavern I had just entered, the system generated a vivid, firelit scene. It paired this with an audio narration that played over 21 seconds. I had not asked for any of these features individually. Yet the system provided them seamlessly. It did so all in one go. It demonstrated a level of automation and integration that was both impressive and unexpected.

A demo that works

Among the many additions to the model list in the update was RPG-HaloTales-V1. This component package was created by AMD itself. This package wasn't a single model. It was a combination of four distinct backends. Qwen3.6-27B handled text. Flux generated images. Whisper listened to audio input. Kokoro produced speech. Together, they created a powerful, multimodal system. It could process and respond across multiple AI functions at once. It loaded into memory almost instantly.

I tested this setup by sending a simple request asking to describe the tavern I had just walked into and to show what I would see. The result wasn't just a paragraph of text, but a fully narrated scene with an oil-painting style image and spoken narration. The system also maintained consistency, keeping track of the room's layout and the appearances of the people introduced. For example, the scarred barkeep remained consistent from the first prompt through the next 12 turns, without any need for rerunning the image generation process.

Routing and privacy

Lemonade Server, now in its updated form, doesn't just serve AI models. It acts as a routing system. It allows users to create complex, multimodal workflows. This is done by defining a Hybrid Router. You can specify a set of models with varying capacities. Let the system choose the best fit for each request. This decision-making process can rely on a small language model. It analyzes the prompt or uses boolean rules set by the user. These are configured using a visual builder that incorporates regex conditions. The system's routing logic is designed to be smart and flexible. There are default behaviors that may need user adjustments. This optimizes performance.

One of the key features of the routing system is its privacy gate functionality. If a prompt contains personal data, such as identifiable information (PII), the system can be configured to send the request to a local model instead of forwarding it to a cloud provider. Conversely, if the prompt doesn't include such information, it can safely be sent to the cloud. However, the default settings for this system are not as safe as they could be. By default, the system is set to 'match_false,' which means if it fails to match a rule, it will send the data to the cloud. This could accidentally lead to private information being processed in the cloud, which contradicts the intended behavior. The correct default should be 'match_true' to ensure data only goes to the cloud when explicitly approved.

3D models in under two minutes

I also tested the image-to-3D pipeline using the tavern scene image that the system had previously generated. Using a fully local setup, I watched as the system processed the 2D image and transformed it into a textured 3D glTF file. The entire process took under two minutes, a speed that I found quite impressive. However, the outcome wasn't exactly what I expected. Instead of producing a full 3D model of the tavern, the system focused on two cloaked figures in the image, separating them from the background and giving them a somewhat three-dimensional appearance.

These figures looked convincing when viewed head-on, but as I changed the viewpoint, the models began to appear flat again. It seems the system had misunderstood the task, perhaps interpreting the request as focusing on the figures rather than the tavern itself. Despite the confusion, the speed and capability of the pipeline were clear indicators of its potential. The system demonstrated that it could handle complex 3D generation quickly, even when the result wasn't exactly as intended.

Frequently asked questions

Can the Ryzen AI Max+ 395 create 3D models?

Yes, it can generate textured 3D models from flat images using the image-to-3D pipeline.

How does the model router work?

The router selects the best model for a request based on prompt analysis or predefined boolean rules.

Based on reporting by XDA Developers, compiled by the Tradingbird newsroom. Published 02 Aug 2026, 15:17.
Topics: AI · Hardware · Software
Read this in: English · Arabiy · Deutsch · Espanol · Italiano · Portugues · Russkij · Turkce