Jia Shen (00:00.046) 'Cause for us we as as an engineer you think, I have a great character engine, we have great animation, we know how to like make a a character like look very, very human and then then you next thing you think is like how do you empower it? Jeff Lu (00:10.808) Core models are mailing our trend, especially for avatars. for avatars we actually develop the whole tech stack. But we also have some other things that's more on the video foundation model side. Richard Bowdler (00:23.766) We sit on top of the the shoulders of L L and Voice AI giants. we we have built the platform so that you can kind of bring your own brain as we're terming it. Jia Shen (00:37.71) It's just like it's like you can have conversational characters, but the question's always gonna be, Why am I talking to this character? Why do I care? Hermes Frangoudis (00:52.248) Hey everyone, welcome to the Convo AI World Podcast, where we talk to the builders and founders shaping the future of conversational AI. Today, we have a very special episode focused on AI avatars and digital humans. We're speaking with the researchers, developers, and founders building the next generation of real time avatars and exploring what happens when conversational AI gets a face, a voice, and a more human way to interact. Let's dive in. What was the core problem or catalyst that led you to found the company? Jeff Lu (01:24.682) So Akool was founded about four years ago and at that time we noticed that a lot of the video generation technology is becoming mature, but we didn't see much products out there on the market. We believe the video generation can help many organizations and teams to do lots of work because creating video is very hard and also want to do video to do interaction is very hard. So that's how we get started and just make video creation much easier, more personalized, make communication easier and also make it more interactive and fun. Hermes Frangoudis (02:03.874) That's so cool. Let let's dive in and kind of unpack that. So with what you guys do at A Cool, you do a lot of visual storytelling, right? And is that correct? That's right Yeah. So when you think about visual storytelling, how does Akool do these things that no one else seems to be able to do? Jeff Lu (02:23.606) Yeah. So on the video story holding, I think a person is the most important thing in the story. So we really focused on the person, the characters in the story and try to generate photorealistic and peoples and so on that it's very hard to tell it's AI or not. People really like that. And meanwhile, very unique stuff is we are doing lots of this stuff in live real time. as well as running on the edge devices rather than depending on the cloud compute. And then that's also changes changes it a lot. So imagine you are able to like interact with live characters, like persons in a video in real time. I think that's kind of different level of experience. Hermes Frangoudis (03:14.252) Yeah, that that face to face interaction really changes how the feeling that you get. Jeff Lu (03:22.21) Yeah, that's right. Hermes Frangoudis (03:23.053) So, you talk about not only running in the cloud, but running on edge. So what does that mean? Like reducing those sort of technic technical barriers? Jeff Lu (03:32.438) Yeah, so now lots of the AI models are pretty big and you need to run them on the cloud with high energy GPUs and the compute. What we do, we are doing a lot of things, optimizing the compute and optimizing the stuff to make them run on their edge devices and especially on the devices such as your laptop, the mobile phones, and so on. So there are several advantages. Derek Zheng (04:01.4) So we know for many kind of platforms, they are pushing for to be more like human beings, right? For the lip sync, for the body gestures, but you are doing differently, right? You are pushing for a Japanese enemy inspired style. So why do you choose a different way? Jia Shen (04:18) There's a lot of reasons for it. Easiest reason to say is that we are in Japan and that like animated characters is what is a widely widely accepted communication format. The real thing is what we've learned from our like the VTuber days. You our our background as AKA is that we are a virtual like VTuber virtual character company from a technology and an agency standpoint. and when we entered the business like five years ago, like for instance, everybody from America to China was investing big in virtual influencers actually. You know, those Comic Kong genies, China was like plying money into it. Korea was s spending a lot of money into it. And everybody was trying to do the same thing. They want to do lifelike realistic virtual celebrities. And none of that worked. Like you can't say a single one's work, right? And it's easy to say for us, like, but the real reason is that the whole point on cartoons is to be a caricature. And so first of all, I would draw you as a cartoon character. Like I don't need a digital clone of you. I need a cartoon character version. I need a specific just like where I was saying you want to do the fortune teller, what's like a very concise experience. If I replicated you as an AI assistant, I don't need you. I need the assistant version of you, right? Both your skill set is as subset of you as well as your visuals should be a subset of you. And I think that's a very important thing to understand. And then so when then you present it to people, they will also basically understand that it's an AI assistant. We all know that assistants that we have aren't replaceable for a human. They can't solve everything. And that's important for expectation setting for what you actually put out there. So for us like virtual characters, the biggest piece for our influencer side is go more caricature, go more cartoon. Make sure that they understand this is Like another another world. And then because of that, they're willing to dive more into it, believe more into it, because now they don't care about the surface. The surface is this beautiful surface, but like it's a cartoon, right? And it actually is much more powerful when you are animated. I I strongly believe in that. Derek Zheng (06:16) Yeah, I agree, I agree. Like a way push for the human like or a left like, it always distracts me like is it real enough? That's what I'm focused on. Jia Shen (06:27.416) Yeah. But if it's a character like for us like carts like what we notice with when we do VTubers, they want to be we want to make them look like anime, right? Mm-hmm. But when we're doing Nascot characters, like we're working on some prefecture characters, they're just like hand drawn. And actually lower frame rate is better. You don't want it to look that animated. You want it I mean like that like high frame rate. The more like it feels more like cartoon, the better it is, right? and and i I I think a lot of people even agree with like, you know, for us like Kray Lo Shinchan, like Lavi Shashi. You know, you just had the C G movie come out. It's okay. Yeah. But the cartoons better. Right. And what you if you were to make an AI character of him, like the C G version of him, obviously it would be okay, but it'd be much more of a good experience where it was just that's what the cartoon character version talking to you. And so we I I I firmly believe in animated Hermes Frangoudis (07:14.774) Yeah. Like what inspired the founding team to focus on like digital avatars in the in this space? Richard Bowdler (07:19.532) Yeah, absolutely. The founders of Trulience, this is not their first radio. They have a f previous exit, the video streaming tech and that they sold to one of the large ed tech platforms out there, called a Blackboard. Well for that, wanted to go on and build something that is still again, very much a a frontier tech. Merrick started building Trilliance it's about five years ago now, kind of pre, you know pre the the explosion of AI. And it was a combination of him wanting to build something that was really on the on the kind of frontier of what was possible with technology. And also on the personal side of things, one thing that sort of drove him to this struck a chord with him was a thought that his he had a an old elderly grandmother who was in a care home quite far away from where he lived. And his thought was, it's just such a shame that she sort of sits there languishing and doesn't have the level of attention or interaction that she could have. I wonder if there's a technical solution to this, as well as a human solution. And it was a seed of an idea from which Trilliance was born. Now that that idea has evolved and the the kind of guiding vision is How can we give AI face and a body and to make it more human? Hermes Frangoudis (08:46.894) It's super interesting. It's like very human touch to that to the story. Like the inspiration is how do I bring that that face to face interaction back? Right. Super cool. Really what what makes Aquel's model architecture, like data training, like what's unique about it? Jeff Lu (09:03.168) Yeah, so on the technology side, we really focus on developing as much technologies in-house as possible. So lots of the things that we do, we try to make it in-house and do it ourselves. And meanwhile, we do leverage some open source models, mainly on the foundation model side. And our core work over there is to make the learning faster. And with more resource constraints and ch fine tune with our data and so on. So we're a pretty tech heavy team and we believe that tech differentiation is very important in the market and we need to keep improving on that. Hermes Frangoudis (09:46.67) Super cool to hear. So you use a mix of combination of foundational models like open source, but also your own models that you're training, right, specifically for this? Jeff Lu (09:57.026) Yeah, the core models are mailing our trend, especially for avatars. for avatars, we actually developed the whole tech stack. But we also have some other things that's more on the video foundation model side, not directly related to the avatar business, but these are mainland built on top of open source and all PMI on top of open source. Derek Zheng (10:19.07) And Ling, we also would like to hear your thoughts about the future of virtual economy. So do you think it is good enough so or it is g still going to increase? Jia Shen (10:30) The virtual economy, like it's such a broad term, right? Yeah. I had a good conversation with the with our Indonesia team today. Our business there is is robust, but Indonesia is not a big money spending country. they don't spend a lot of money, like and so like the people are always like, what's your like? You know, average revenue per user, that kind of stuff, right? And the truth of it is, is that they're not understanding how to approach the common person. The funny thing, the thing that I point up to is that, you know, TikTok EC e commerce, the basic the there's the online store, their second largest country in the entire world, I guess third, I guess, outside of it's China, America, Indonesia. Really? Indonesia's huge. Indonesia ranks at the same level as the United States as far as revenue annually. It's right. And it's a pure volume thing. They just have a lot of young people that want to spend on very small affordable things. And like you can buy a lot of small affordable things, I guess, like cards and like whatever. But in the end, the best small affordable thing that they should be buying is virtual. Right. And so yeah, I think virtual economy I can answer this question like three thousand different ways. The way to do things digitally is much better than the offline way. And I think when we originally entered the virtual character space, it was a very important realization. Young people don't care if your influencer is a real person or not. They just want to believe in the value and that it's they believe that it's real to them. And I think that's the same for any product that they want to buy. Right? so if you have the proper virtual product, virtual experience, whatever it is, like, yeah, they'd rather do it virtually. Like this world is, you know, changing so drastically. And I think that's you know, I I thought that, you know, five years ago, but this even in the this month it's actually been very, very obvious to me. You know, you look at all the video AEI generation, it's good. It's too good. You can't tell. Like you can see something like where like I can't tell if this is real or not. Did this really happen? Is this the real news? And so what's gonna happen is anything that looks pretty real, nobody's gonna believe it. Right? Like so the concept of what people value, what they perceive now, is gonna be fundamentally changed. It's really gonna change. There's no before it's like, I filmed you doing something crazy in the street. And you're like, I can't believe you did it. Now you're like, Yeah, I don't know if you did that. You no longer care about that video from your iPhone anymore because you're just like That could have been faked. Yeah. It's at a level where it's so not believable. You just have to like basically believe in kind of a core system that you fund them but they care about, right? So entertainment systems, other type of virtual like like like like subs sis systems, those are all the important pieces. The next generation's concept of what's real, what's valuable is gonna be very, very different. Hermes Frangoudis (13:00.844) So in in building that, like what were some of those initial challenges maybe in in developing this sort of avatar technology? Richard Bowdler (13:07.714) The world we live in now is very different to what things looked like when we were starting out. When we're starting out it was heavy studio builds, heavy lifts, full three sixty cages of, you know, surround cameras and intensive filming days to make sure that we could capture people and get, you know, lifelike imagery from as many different angles as possible. So there's a a kind of a a huge front end lift. in in building even just fairly rudimentary avatars. So that was a big challenge. And then the other piece, which again the world we live in now is is very different, was trying to get the avatars to be meaningfully and reasonably conversational. I mean now with the kind of explosion of L LMs and the the voice AI layers that sits on top, we we we're kind of sport for choice really in terms of how we how we get that to work, which we're now in a in a world where you can have conversations with AI and it and it feels incredibly fluid, but way back then it was it was tricky and and it's it's taken a while to to get to this point. Hermes Frangoudis (14:17.996) Man, I couldn't even imagine like the early days of training the AI to understand from all angles at the same time. Like the I remember the the capture cages 'cause we came from the AR background. So we used a lot of that for bringing like virtual avatars into the scene and animating them and boy were those things heavy. So it's super cool that that's kind of the roots of this. And Richard Bowdler (14:36.398) It's kind of evolved, right? So the tech, you know, has evolved through these capture cages and then when I joined, we were doing mo cap stuff and capture suits with dots on, black suits with dots on, and filming that so we could then dress sort of digital mannequins. When we started out, we we were working in Unreal Engine, gaming engine. And you know, building building with CGI. up until I mean we still we no longer build in with gaming engines, but we do still have some CGI avatars because they s fit with certain use cases, certainly in gaming and and so on, where there's less of drive or or a want for hyper hyper real. But the techniques have evolved so much, even just in the last six months. So this side of Christmas, it's kind of things have very much exploded in both appetite and I think it's probably because of where the product is at. And not just Trilliants, there are others out there in the market, but certainly the techniques that we've developed. Hermes Frangoudis (15:38.23) see in that maturity it's definitely like created a new interest and vigoration probably in the market. So when we're talking about the the stack and you say you you've built a lot of this in-house, like what tool in your stack is maybe the hardest to scale or technically in terms of adoption? Like what is one of the maybe pain points you felt along the way? Jeff Lu (16:02.944) Yeah, so the the pain point we felt along the way like mailing around how to create video quality that's that's super high and how to get very precise control and also how to like reduce the cost. So these are some of the problems that we have been constantly working on. And on the cost side, definitely that's very important, make us competitive and makes the real-time device possible. And on the on the other side, the result quality is very important as well. A lot of the customers, especially B2P customers, they expect very good result quality to be able to use them and they are less price sensitive on the B2B side. Hermes Frangoudis (17:01.72) That kind of leads me into my next question. Like, what kind of trade offs do you face between speed, realism, and controllability in the in the generation, right? Because it seems like it's speed versus price sensitivity, but are there other levers that you have to kind of twist and turn? Yeah. Jeff Lu (17:19.0) So there are several factors that they are associated, and we need to twist and make it into the ideal status. The the first piece is actually around the what I would say, definitely the result quality. So most of the customers expect the best result quality. And then the next piece is how to run it faster, more efficiently. reduce the cost, increase the speed, and all these kind of things. Then the next piece is ease of use. The if your tool is too complex, it's very hard to use and people will get into confusions and so on. But if your tool is too simple and don't give user too much freedom, then the professional users will have lots of difficulties about all these kind of things. And also the flexibility we are balancing, like how much flexibility we want to give to the user in terms of their APIs and so on, as well as back to the how easy it is to use. So quite many factors are combined and there are lots of balancing. Derek Zheng (18:31.758) Okay, okay. So you you also mentioned the DI, right? DI platform. So several times. So can you explain to our audience what it is and how it serves the framework of building AI Avatar? Jia Shen (18:46) Yeah, so our company, like, you know, the the founders like me and my brother were both technology background, right? And so we built companies based off of technology first. Mm-hmm. so the fundamental base for AK virtual is an animation engine for for characters and that that allows us to do, you know, generally whatever we want, but we had applied it to VTubers animation. And that's why we actually started looking at AI to power these characters pretty early on, like basically like three years ago. because for us we as as an engineer, you think, I have great character engine. We have great animation. We know how to like make a a character like look very, very human. And then then you next thing you think is like how do you empower it, right? Like how do give it decision making? How do you give it personality? How do you give it intelligence? So DI is kind of a product manifestation of that. It allows us internally initially to build out, like I said, the two different types of AI concierge and assistants. One is very functional, which is to do a job in this case, which is like information in a shopping mall or an airport, or navigation through Shubia or a train station. The other one is basically basically checkout counter, self self register. And so you basically helping use go with like in a convenience store, where is the milk? What's where's this thing and that c that thing and then helping you walk through the sales process. And that and then another key benefit is that all of our characters are like multilingual. Japan right now it's like what is it, forty five million tourists this year? Right, which is like if you think about it, Japan has a hundred and thirty million people and like forty five million tourists every year is like there's a there's a language problem too. So that's one basically functional, vocational, AI assistance. The other one is entertainment ones, right? And this is where You know, like for instance, the fortune teller is usually like a mascot. It's still about information, but they have to have more personality. And what they need to basically communicate is more the cultural information, right? So Izum Gutaisha is literally a temple. Obviously the fortune telling thing is fun, but you need to also tell about top Buddhism, you need to talk about why the temple exists, why you want to do this type of fortune telling, why this temple is specifically relevant for this one versus this other. And so that aspect of it is a different expertise, right? Like Because the vocational one is like some personality and mostly a rag database, but the the culture one is not, right? Because like you have to make sure that it really understands the cultural stuff, it can understand it pro and explain it properly. So those are the two pillars of AI. And so the first thing that everybody asks for is like, how can I build my own character? Do you can do an avatar builder like that? And so yeah, our core platform that we will expose at some point is you have to bring your own character or you have to use some basic framework to bring your visual character. But then the rest of it works. And so we we follow friends with the the VRM like three D character like a standard. And so if you bring a V VRM, you can use our platform. If you use live two D for your character for the two D version of it, you can use our platform. and then the cool part about the back end side of it is that it is infrastructure agnostic. Right, and so that's why Agora is like super interesting and important for us, right? Because we do do our own LMs, but the infrastructure piece of it, like and we know new models both from voice and for from intelligence standpoint is coming on all online all the time. We don't intend to lop our platform to one like model or or other. 'Cause we change it ourselves all the time. And so we're we have an a little L L called Shisa and it's actually like basically the strongest Japanese language model. But the it's not a bit it's not a model that can rival GPT like as far as like internet research and all this kind of stuff, right? And so for us since like one of our more advanced agents, good looking character, do all this stuff, the first layer it hits is actually a Shisa. 'Cause it's fast, it's really fast. It's it'll it'll come back in like like like sub s subsecond while it waits for GPT to think, right? And GPT likes like two seconds, three seconds, ten seconds, whatever. Small the smaller AI buys you time. You're like, let me see and it basically does the whole like and conversation things. And so this is where it's important to be smart about how you mix and match what your AI infrastructure is 'cause it's just changing all the time. Hermes Frangoudis (22:43.192) What are some of these key components that enable Trulians' like real time avatars to seem so interactive? Richard Bowdler (22:49.762) Well, you have some of the guts of it. Let me just break it down in simple terms. Trilliant, we deliver the the visual front end. We give AI a face and a body where needs be. And that's both hyper real humans, all characters. So, you know, like a an off trademark Mickey Mouse or like a fun, fluffy dog. So we are that. We sit on top of the the shoulders of LLM and voice AI giants. We have built the platform so that you can kind of bring your own brain as we're terming it. So you can plug in any language model, any voice AI tech, and we we play around with different combinations of those to see what is it that really plays nicely together, like which language models, where are we hosting them, privately, publicly, etc. and which versions of those models are best fit. And then both on the speech to text and text to speech, we Play around with with those pieces too. And by fusing those together, we can then have a real time, realistic latency, a conversational interaction with an avatar. It's worth pointing out a kind of curious quirk here. Now, there is voice-to-voice tech, and we have integrated voice-to-voice tech, but the latency on that is we you can dial the the latency kind of up and down. There's a threshold of about five hundred milliseconds response time. So from when I speak, then there's a gap and you speak and natural interrupts. Anything below about five hundred milliseconds, like you ask a thing a question, if it's below five hundred milliseconds, it to a human it feels like it's happening too quickly. Yeah, it feels weirdly too quick. We have got two To the point where we are as fast as we need to be, faster than we need to be. So it's kind of quite an interesting quirk. One of the technical difficulties is in lip syncing. So having it's called the the visimes, the shapes that lips make with different people speaking different sounds. And then you multiply that across into different languages. And then you have all different other shapes that people make with their mouths when you say. perhaps speak in Arabic or Hindi or what have you. Because it's a massive technical challenge, but we have a a team that they they focus exclusively on lip syncing and mouth shapes. Hermes Frangoudis (25:21.254) In terms of what you're getting out of it, what role does like real-time inference play in products like your streaming avatars and live camera? Jeff Lu (25:29.378) Yeah, so in streaming avatar and so on, there are quite many things ongoing over there. So we do quite many things around AI agents and AI agents related. We integrate into many like AI agents platforms and also with hardware as well to make make things happen. Also translation is also a very interesting one for translating meetings, no other things, in real time, live and so on. Hermes Frangoudis (26:05.944) Yeah, so you can have two options, right? Like you could be live translating a person and then live or communicating in real time with a completely generative avatar. Yeah. But those are both wildly different directions, it feels like, but under the hood technologically, is it like basically solving the same problem? Or is there different nuances in each one and each approach? Jeff Lu (26:34.318) I think they are many of the technologies, the underlying technologies are connected. So it's it's it's just about how they are being tuned and adjusted, how the system is designed and so on. But lots of technology we are talking about, the underlying framework is similar. So we definitely try to develop. new solutions based on similar foundations. Otherwise the work load on our side will be extremely high as well. So kind of expand the feature family group, which is interconnected. So when we launch new features, we don't need to write everything from scratch. We can leverage our existing things to amplify our product features. Nice. Hermes Frangoudis (27:23.286) It's awesome. How everything kind of like works to build on itself and you never really like starting over just from scratch again. Can you walk us through a little bit about what makes the the video translation more advanced than like your typical dubbing or like lip sync style technology? Yeah. Jeff Lu (27:43.094) Yeah, so there are several things that we do in the video translation. First is we clone the voice and we translate the language and we make sure it fits a video lens and with the portal 150 plus different language. And the second is we do the whole face reanimation rather than just the mouse region. So we do a lot of reanimations of the people to face and so on. And also we have a very well-designed workflow that allows people to easily get input a video and get everything clear and then output the results and so on. So we also have lots of freedom to allow the users to post adding their results to make sure the translation or the things are accurate and the kind of stuff. And what's more, we are the only one that's able to put the whole pipeline to live real time, which means real-time interactions with translations and then that happens in all the different places, including I think even webinar and meetings in the world in different places. So definitely that's a lot. Hermes Frangoudis (28:59.598) That's huge for like those massive global town hall. Now you can have your senior leadership speaking in their native languages, but having it come to the audience in their own native language with all the proper reanimation of the face to make it feel realistic and not just like it's a cheap dub. Jeff Lu (29:21.068) Yeah, yeah, that's right. Hermes Frangoudis (29:22.338) So how does your system handle like maybe multilingual conversations or like real time language detection? Is that passed through to the L LMs and then it's more just so about like syncing of the motion? Richard Bowdler (29:32.696) Exactly that. We don't handle that. we we work with that. You know, there are obviously lots of providers out there that do that. So you don't Hermes Frangoudis (29:40.598) Don't really constrain the customer in that sense. You let them bring their super cool Richard Bowdler (29:44.692) Exactly. And there are some so Azure is pretty good for multilingual within the context of one one conversation. So an avatar can listen in Arabic and if asked to can respond in Spanish or or or whatever combination of languages from their their selection. It's kind of awesome to experience. You're like, wow, that's crazy. Hermes Frangoudis (30:10.508) probably for like someone that is multilingual probably feels actually very like at home because like I speak Greek and English, right? And with my parents, it just depends on what words like the simpler one to say. So you get like this weird like mix of the two languages. I w it'd be curious to like play around and and see how it works with that. Derek Zheng (30:28.002) So we see the concert, but we want to learn from technical perspective. Singing technically is it harder than just normally speaking? Or is it just that the sing Jia Shen (30:39) You mean AI singing? Derek Zheng (30:40) Yeah, AI singing. Jia Shen (30:41) Well we don't do AI singing for our concerts, but the the AI singing is I mean in real time, no. But like I mean right at this point AI singing generated is very good. It's the same thing as songwriting. You just need to make sure you write specifically for the song that you want. Like the yeah, I'm not sure you're actual question 'cause like the from a voice model standpoint. That's my question. From a voice model perspective, yeah. Yeah, voice and music models right now are great. They're very good. they're a little slow, and that's our bigger problem. Like and we focus mostly on the speaking part because you know, we want our AI guys and girls to feel real. Mm-hmm. Right. And so the Western models, like, you know, eleven labs or whatever, they're really good, except the Japanese is terrible. It's just really not good. And then the the models that are good at AI a Japanese sound like are usually the Chinese models. You know, we we have our own model too, that we think is pretty good. but, you know, it's always in it's like an instrument. Right? Like it's like they're good at certain types of voices, but other types of voices no, right? Like and so because we're in the business of creating characters, yes. It's not just like, this is the best male Japanese voice, great, that's okay. But I need to have this anime girl or this like this like famous character that's really really pitchy or you're re replicating the celebrity voice that you have to really make sure you hit the the comedic nuances of the voice and that's a lot harder. And that requires more really more programming on our part, like I'd say. We're very we know all the details about it. And yeah, this is this is why we got up to this part of L AI is because entertainment version of Japanese is different than the core version of Japanese where you know, people just don't speak like anime characters all the time. But if you're talking to anime character or like this kind of like even Whatever, like they role play the voice and the even the words that they use. The it's a lot of fun and that's where a lot of specialization can happen. Hermes Frangoudis (32:33.568) In terms of giving it a face, like one of the really important pieces of face and conversation is the emotional nuance and naturalness of the the interaction. How does Trulliance really work in that sense and and try to capture that piece of the magic? Richard Bowdler (32:49.164) Yeah, so our two main categories of avatar types are one is an image based avatar and one is a video based avatar. So the image based avatars, you can take you could take a screenshot of the of you or me and we could then turn that into an avatar. Or we could take take a video section of us speaking and turn that that video into an avatar. And both would give, you know, very realistic lifelike representations. What we then do is we have a range of emotions that we have baked into our our model and we will then map those to, you know, a sur surprised response, ra raising of eyebrows, furrowing of brow, and so on. There are things that we can do with that. It's it is one of the areas like with lip syncing that is tough to get really good because there's a lot that needs to happen to kind of change a face to to represent emotion and also to then synchronize that with whatever is the the kind of appropriate emotion to display. but yeah there's it's a it's an ongoing bit of work. Hermes Frangoudis (33:57.612) No, it makes sense. It's like it's one of those really challenging pieces 'cause everyone's face moves a little bit differently and and there's a lot of nuance into that. But getting it close and from what I've seen on your avatars are really awesome. Richard Bowdler (34:11.34) Yeah, I mean also just to say, you know, get over a bit of the tease, what I'll do and it would be cool if I'll share with you a link to a repository of avatars that people can go and play with. Hermes Frangoudis (34:22.734) That'd be epic. Richard Bowdler (34:23) It's awesome. And that's a combo of I think it is mostly I'll make sure that there's a a bit of a mix of both video based avatars. I'm I'm one of them. There's a phone of me. It's me who sat in my home downstairs. And I've done some fun experiment like you know, content experiments where I've done like recorded me talking to AI me, which is it's always a slightly creepy experience. Hermes Frangoudis (34:49.676) Very uncanny valley. Hermes Frangoudis (34:53.484) Yeah, they're like, oh dude what Hermes Frangoudis (34:56.247) Speaking of customers, how do ACOL's customers use the tools today? And like, what do you really see as some of the most common use cases versus maybe some of the more surprising? Jeff Lu (35:05.184) Yeah, yeah. So for the use case and so on, that's there are quite many different use cases. But we start with marketing advertisements and then we go to more of the things like film productions and we go more expanded into the field such as powering other AI agent companies with a face and a lot of the combinate internal communications. So yeah, these are the common use cases that we Hermes Frangoudis (35:39.034) I think you touched upon some like really interesting use cases. I'd love to dive a little bit into each of those. So you said you started marketing, advertising. I think that kind of makes sense, right? Like they are the hungriest for content and probably the most strapped for budget when it comes to creating big content. But you you said you guys also do internal communications. So can you tell me a little bit about that kind of use case? Jeff Lu (36:06.28) Yeah. in terms of that, it's really related to translation. So we are able to translate videos into multiple different languages from either static video or live video. And we can do that in 150 plus different languages. So for many large organizations, their leadership team might only be able to speak one or two languages. But we are able to translate their speech into like tens of different languages and distribute to their employees globally. So it applies to a lot of the live meetings as well and so on. Hermes Frangoudis (36:49.846) In terms of the the customers, right? Like what feedback have you been hearing in terms of like the effectiveness of the avatars and these sort of real world applications like Richard Bowdler (37:00.16) Yeah, yeah. There's I love talking to customers. So I was asked earlier today, like where do where do I spend most of my time? I spend most of my time f I it's across probably three areas, but most of the time talking to customers and figuring out like how do we get the product and exactly to work for for for them. Working with partners. so like yourselves, like LLM, you know, big L companies voice providers and then getting on podcasts and and talking to people like this because I I I love doing this. It's really interesting with with speaking to customers. I'll give you an example. So one of our one of our case studies is working with a government in India, a state government, and they wanted to provide access to healthcare to a illiterate population. What we did, we worked with a a local healthcare authority to set up screens around around a village so that the illiterate population could then walk up to a screen and speak to an avatar in Hindi and have a real interaction and gain access to healthcare knowledge and healthcare understanding and be signposted to how to go and seek appropriate healthcare. So the kind of feedback that we get is, I mean, like it's a transformative technology, right? When you can literally transform someone's situation from not being able to access something to then being able to access it and transform whatever is going on with them from a healthcare perspective. I mean that's that's just huge. Hermes Frangoudis (38:51.118) That sounds epic. Like the fact that people no longer have to rely on their ability to read and write in order to access knowledge, right? Because that that's been a significant like barrier of entry. And we know across the world there's significant piece of the population that just aren't in that position, right? Like they they can speak, they they're able to communicate in that sense, but the world just shuts down when you talk about books and writing. Richard Bowdler (39:14.688) Exactly. And this, you know, having natural language be the interface is a game changer. And then having an avatar visual front end as the kind of the compelling Hermes Frangoudis (39:25.582) Humanize it, right? Richard Bowdler (39:44.246) humanize that, absolutely. And to draw draw people to that as well as have people feel comfortable interacting, I mean it's a massive deal. I mean, when we went into this this project, I started digging into the figures about illiteracy globally. There guess you you You would not think it, but g have a guess of how many people globally how many adults globally are Hermes Frangoudis (39:50.35) I don't know, Gonna need your help on in terms of super surprising or just higher than you'd probably expect. Richard Bowdler (39:57.218) Slightly higher than you might think. Hermes Frangoudis (39:59.992) Twenty five percent. Mm-hmm. Richard Bowdler (40:02.232) Not Far off actually, of the world. I mean Yeah. eight hundred million. Pine kind of directionally, you know, getting up to a billion of of an the adult global population. So maybe maybe that is about a quarter of the adult of the global adult population. Astonishing, right? I'm not suggesting that just having interactive avatars is a panacea to illiteracy. I think, you know, if Hermes Frangoudis (40:27.798) But it can open up access to so many things that were traditionally inaccessible or people had to bring someone else to help them through this process, they can now do it themselves in the same way that they communicate with the people around. Richard Bowdler (40:39.784) Exactly. Exactly. Hermes Frangoudis (40:41.656) And I feel like that's just the power of conversational AI in general. And then you're putting a face to it, which makes you feel even more comfortable, like you're talking to something, not just like a dot on the screen that's like visualizing your sound waves, right? Richard Bowdler (40:55.766) Yeah, yeah, Hermes Frangoudis (40:57.1) It kind of moving into the the competitive side of things. Like when we first started, you you you mentioned there's some other Avatar players in the space. How does Trillians really differentiate itself and like what what do you guys do to stand out? Well Richard Bowdler (41:09.902) I'll tell you I'll I'll tell you what it is and I'll also tell you a bit of the backstory to it as well. So standout differentiator between us, there are a handful of, you know, comparable in terms of look and feel avatar platforms out there. They, to my knowledge, none of them render avatars how we do. So we render everything client side. So it's all we're essentially leveraging the compute power of the end user. We render the avatars in the browser in in simple terms. Now, what that means is it's pretty cheap for us to do that. compared to if we're needing to spin up a new server for every single instance. Now, there are use cases where it doesn't matter, right? So we have projects where we've had avatars in those sort of hologram holo boxes in you know, shopping malls as a kind of meet and greet, you know, direction finding person. Hermes Frangoudis (42:05.07) So with the core, it's really about delivering upon the mission and and understanding this is very clearly where we're going. But for the other pieces, it's hearing the customer needs, hearing the maturity of the market opportunity and deciding which makes sense to bring in. Do you bring it into the core, or does it still kind of stay as this peripheral thing that only gets love based on how much market opportunity there Jeff Lu (42:32.274) Yeah, we run things a little bit like labs and so on. So for the for the new features we think there might be opportunity. We kind of launch a lightweighted first and see how much tractions we get. We value like user feedback and so on. And if we decide that's getting us a huge amount of tractions and so on, then we will move to getting that into our co offering and increase their priorities and so on. Hermes Frangoudis (43:04.288) Makes sense. So you you roll it out almost like like beta, test the water, see how much interest there is in the market, right? Jeff Lu (43:11.18) Yeah, yeah, that's right. Hermes Frangoudis (43:12.18) In that sense, like how do you balance experimentation with stability? Like when you roll out a new feature, it is it fairly early, or do you try to make it mature enough that when you roll it out, it's not gonna cause issues? Jeff Lu (43:26.382) we wait until the feature is mature, then we roll out. before that we might put it in beta and people can test out or we might have some invited parties to come and use them. But for our official launch they are all mature product. So just make sure that people have a consistent expectation of the product and qualities of the overall platform. Derek Zheng (43:54.626) I noticed at the TGS you showcased the two fascinating demos, right? AI Witch and the Fortune Tale. So how's the feedback from the visitors? Jia Shen (44:03.046) obviously I'll say the feedback was great, but the feedback was really good. The g w we spent a lot of time thinking about the the the KPIs of attracting people to come, getting people to interact and having them feel like they came walked away with something, right? Like that's I'd like to say the core challenge of everybody that's attending today's conference, right? For it's just like it's like You can have conversational characters, but the question is always going to be, why am I talking to this character? Why do I care? and do I walk away with kind of a fun experience? So Tokyo Game Show is show culmination of a bunch of things that we do, but it's it's I think it's more interesting to highlight what we have been doing that's allowed us to be able to choose the experiences that we did for Tokyo Game Show. For you know, we have a product called Day Eye. And Day is actually two products. One product is an information guidance AI assistant. And so it's like a person in a suit speaking not entertainment and basically providing information. That's something that usually people know why they're going to the character, and the character knows exactly what their job is. And that's a very straightforward AI solution. You should do it well and then it should be very smart about navigation, communication, and that's it. And it should be stupid about everything else. And entertainment character is totally different. And so putting it for a talking game show, well we did We have three characters I think that are very interesting that are all in the wild. We did one for Osaka Expo. and that was for Yamanashi with kind of a cool b a big agency here. we did one for this kind of game center out in Yokohama called Hapipi Land. And then we did this one with this temple called Izumo Taisha, it's famous. Like the Kyoto Temple. and they were s actually they did this big good goods pop up in Motasando. And that one was our most interesting one because that one was a fortune telling fortune telling character. With an entertainment character, you need to make sure that they have a very specific reason to exist. And so when you're normal or we're thinking about Japanese people or tourists, these at this point, see a character, a child a child will walk up to a character like, this is very entertaining. What's it doing? it's gonna bunch of buttons and try a bunch of stuff. An adult will be like, they don't even know what to say. Right? Like it's this whole this whole prompting thing. And so with a fortune teller, the experience is very expected. And it's actually can be very, very good and very specialized. And then the user has something to take away. You walk up and the character the the character just like, you know, what's your zodiac sign? What year were you born? Mm-hmm. you know, like, asks you some questions and you're you expect that, you answer it, and they will role play, like for instance in the case of Izumanatai Sha, like a specific spirit. And in the case of like, you know, for our witch and fortune teller demo for Tokyo Game Show is basically a witch or actually one of them is like one of our handsome guys from other one of our Erix characters. And so that's all actually now a fun experience. It's like a mini game and when they're done with it, they get to take away a fortune. Omikuji. And so they that's really so the whole experience is actually very good. Otherwise you can have an AI character who's like, Kai, what's the weather? What do you think about what's happening in global politics? So like whatever. That experience is terrible. So for the the design between an AI character's user interface is very, very important and that's what we spend most of our time on. Cool. Hermes Frangoudis (47:10.742) Speaking of like implementations and use cases of it, what industries or applications have you guys seen the most impact from from these sort of technologies? Richard Bowdler (47:17.746) Yeah. I mean, we are industry agnostic, it's worth saying, as a technology. people build businesses on top of our tech. So they take it, they run in their own direction in their own vertical. And that's cool. That that is exactly what we want to enable. Where we see a lot of interest and application, so I mean, broadly, as with any kind of chat functionality, sits underneath an umbrella of customer service agent, right? But then if you kind of start slicing into that. That looks like in healthcare, patient aftercare, customer support, kind of frontline customer support in e-commerce, both as a sales agent and and genuine customer support. And then a prod like in in the in the world of product, we have this baked into our platform, which you'll see if you can kind of sign up. We have an avatar that is an onboarding agent inside the product. So you go into the avatar creator and you have a An avatar saying, hey, you're now in the avatar creator, blah, blah, blah this is how you need to create an avatar. Slightly meta, but we're we're seeing quite a lot of interest and application where historically in in e-commerce and in product onboarding, where you may have a chat function in the bottom, whatever bottom right corner of a website, people want and see value in there being, hey, here's here's a a visual representation of where previously I just had some text input. Then also in where I'm because of a lot of my background in learning and personal development, is in the learning and development space. So experts, coaches, advisors wanting to use our technology to leverage themselves and kind of clone clone themselves, which you you can do with a knowledge base and custom GPTs, this kind of thing, but then also take it that one step further and create a a visual clone of themselves that their customer base, their advisor base can have access to twenty four seven. For me, that is a like I mean it the appetite is huge there. And the industry is just vast. Hermes Frangoudis (49:24.418) Yeah, 'cause when you think about it, these are people that have figured out how to successfully build a business around themselves, right? And the biggest constraint is probably their time and their ability to meet with these people because to train someone else to do this exact same thing is sometimes a risk, sometimes just not doable. There's not the right people. So being able to have that clone of yourself, I feel like is huge for those sort of creators. Richard Bowdler (49:46.732) Yeah, I I think it opens up kind of quite an interesting space, right? Where you are both it's like in this sort of information and advisor expert space. Like there there seems to be for a lot of people of a a a a barrier or a fear to you know, they feel like I mean we were talking about this before we kind of came online, came on air, giving away the secret source. It's like, man, give away the secret source. Give people access to the secret source. People will still want to pay you to Go and implement the thing because they don't have time when we're talking about the open source community and the business models that sit on top of that. I kind of see the same thing playing out, but in the information expert space. It's like, man, open source your stuff. Like give people access to you. Like twenty four seven, create, you know, AI MS. Like, go. What that what that does is democratize access to expertise and you know We're we're already seeing it with all of the language models, a sort of a democratization or a like an infrastructure layer of access to intelligence. I think it it it sits on top of that and it kind of slices into the verticals. It's like democratize and, you know, give like give open access to expertise. so yeah, it's a it's a really exciting it's huge. It's really exciting time.