AGI PLAYGROUND

Insta360 Founder JK Liu: The Camera Agent Battle Is Beyond the Lens

The imaging devices of the future should be light at the front and heavy at the back.

Share: Share on X LinkedIn

That's how Insta360 founder JK Liu framed the future of the imaging industry at AGI Playground 2026. Last month, Insta360 unveiled its latest product vision: building a "Cameraman" - an imaging experience defined by simpler operation and better delivery.

"You could also think of it as a Camera Agent," Liu said.

The product will likely take many forms. Even with a light front end, different shooting scenarios still call for different lenses, different focal lengths, and different formats - first-person, third-person, and so on.

As in embodied AI, letting the device itself move freely through general environments remains a real challenge. But a Cameraman may not need a general-purpose brain. When it comes to recording and sharing life, taste that stays on point matters far more than intelligently executing standardized tasks.

Insta360 is already offering cloud-based AI editing directly to users, Liu said. It costs 6 RMB per clip. In the first two months of its limited rollout, the service has produced hundreds of thousands of edits, with an export rate - the imaging equivalent of the "roll rate" in AI video generation - above 50%.

But it's still early. A Camera Agent isn't just a standalone entity or service. What it has to improve isn't only its own intelligence and taste, but a deeper understanding of what people expect, aesthetically, from the physical world.

What follows is the conversation between Insta360 founder JK Liu and GeekPark Founder and CEO Jack Zhang (Zhang Peng).


01 · The Cameraman: the core isn't shooting, it's delivering the result

Jack Zhang: Welcome to AGI Playground, JK. Insta360 recently updated its product mission. Where you used to give consumers a "camera," now you want to give them a "Cameraman." What does that shift mean?

JK Liu: The Cameraman idea didn't arrive all at once. It started back in 2019. We mounted an action camera we were about to launch onto an FPV drone to shoot an ad, and the footage came out beautifully. But capturing shots like that takes real skill to fly and operate. So we worked backward from there: how do we let anyone capture footage like this more easily?

Push that further and you have to see the need behind the need. What people often really want is to stay immersed in the moment, while still coming away with photos and videos that capture and preserve those meaningful experiences. When people go out today, some of them hire a photographer to follow them around. This event has cameramen catching candid shots for us, and we don't have to lift a finger to end up with plenty of good pictures. That's the ceiling of the imaging experience.

Jack Zhang: So from that point on you started taking it apart - what technical means are needed to realize the "cameraman" vision?

JK Liu: To build the future cameraman, a camera needs to do three things well: see, understand, and act.

First, it needs to see more of the world. That’s one reason 360 imaging matters so much to us. A traditional camera captures only a narrow frame, while a 360 camera captures the full scene. That gives both the user and the AI much more context, reduces the risk of missing important moments, and creates a stronger foundation for tracking, reframing, and scene understanding.

Second, it needs spatial intelligence. The camera has to understand depth, motion, orientation, and what’s happening in the environment - not just record pixels, but actually interpret the scene.

Third, it needs to act in real time. That means tracking subjects, stabilizing footage, composing shots, predicting movement, and eventually making smarter filming decisions automatically based on the user’s intent. If it can’t respond instantly, it doesn’t really behave like a cameraman.

So for us, the future cameraman sits at the intersection of sensing, on-device AI, and computational imaging. The goal is for the camera to stop being just a recording tool and start becoming an active creative partner.

Jack Zhang: So the Cameraman is really a direction you'd already thought through internally - you're just sounding the charge now.

JK Liu: Right. We actually talk about it at the company all-hands every year.

Jack Zhang: You just didn't tell us.

JK Liu: We kept piecing the parts together, and now we're going for it. Something clicked when we were talking this morning: if "Cameraman" is expensive to explain, there might be a better word today - Camera Agent. Everyone's talking about agents now, right?

It's probably easier to grasp, without the detour. The core value isn't the process of operating a camera. It's the ability to deliver good photos and videos on its own.

Jack Zhang: We're not buying a gadget, we're using an agent. We've already gotten used to handing tasks to agents in the digital world and getting results back. In the future we'd like imaging to work the same way.

JK Liu: We hope it can work that way in the physical world too.


02 · Light at the front, heavy at the back: the traditional path is moving toward AI

Jack Zhang: Let's start with the imaging line. With the progress in large models, cloud, and on-device compute these past few years, what has changed most about imaging? It's the front-most part of the whole Cameraman. Where does it stand today, and what will the landscape look like?

JK Liu: There really are a lot of new technical paths here. Let me give a few representative examples.

Take denoising. The old approach was traditional CV denoising running from the sensor through the ISP (image signal processor). Starting in 2019, we did end-to-end denoising with AI models, which let even an action camera shoot night scenes well.

Later, the industry started doing tone mapping end-to-end, so in the last few years everyone - us and our peers - made a big leap in color.

But beyond the basics of color, dynamic range, and noise, imaging also has a lot to do with color grading.

That's where you get the things people intuitively think of as "photoshopping" - adding bloom, adding vignettes, layering the image. What we eventually found is that processing this directly through models produces many effects you simply can't achieve on-device.

And that's just for photos. In video, there are far more effects that on-device hardware can't produce at all.

So we think imaging may be going through a paradigm shift in how cameras are built: relying less on the front end. In the future the camera portion at the front could be very light, even very cheap, and it doesn't necessarily have to be premium hardware.

Jack Zhang: The camera is becoming more of a sensor. The optics still matter, but they're no longer the only thing that defines the device

JK Liu: There will always be people who chase traditional craft, of course. But I believe that for everyday users, in most cases, using something light enough and getting a great-looking image matters more.

Jack Zhang: Where does the industry still need to shore things up?

JK Liu: For example, the fine detail you capture with a super-telephoto lens is still very hard for AI to fill in. And cloud generation is asynchronous right now, whereas for a camera, instant capture is a very basic expectation.

So once the models mature, some of this may run back onto the device. On-device compute is currently limited by chip process nodes; once the chips advance, on-device inference and generation should get much stronger.

Either way, one thing hasn't changed: we are seeing more AI-empowered imaging in photography

Jack Zhang: That's the "light at the front, heavy at the back" idea - the optics end still has to meet a standard you can't lower, but the future gains sit further back, tied to AI and to compute.

JK Liu: Right. Take background blur. Traditional optics do this very well, but they basically can't be compressed: how big your sensor is and how wide your aperture is directly determine whether the depth of field and the blur look natural. Simulating that with an algorithm mostly doesn't work. At the same time, AI generation provides another way to achieve a similarly natural result.

Jack Zhang: So if on-device compute keeps climbing, you press the shutter and it's out in a second, and there's actually a generated component inside - is that a race between edge and cloud? Or will something else eventually decide the balance between the two?

JK Liu: I think plenty of industries beyond ours will develop this same branch. Some people are willing to pay less for hardware up front and subscribe over the long run; others would rather buy the service and compute outright, once.

I don't see it as a race, but as two consumption models that will coexist. Some people don't want to pay a lot up front and prefer to pay over time - and for some, this may not even be something they'll use forever, so the cloud is simply cheaper for them. But for high-frequency, must-have users, they may just buy the compute outright.


03 · AI will become a consumer of imaging too - and that's not necessarily a bad thing

Jack Zhang: Would you explore anything at the action layer? When you talk about the Cameraman, have you thought about it having "man"-like ability not just in how it generates and optimizes footage afterward, but at the level of shooting, moving, acting?

JK Liu: The cameraman will have a "brain" of its own - what we call 360-degree awareness of the environment. To capture a good shot or a good video, camera movement is critical: understanding what the subject is, what the scene is, what space or path is traversable, and how to trace that trajectory. That involves the body's own motion, which is fairly complex.

There's also a purely data-level notion of "movement": feed in a wide-angle image and find the good crop.

Jack Zhang: So the Cameraman doesn't necessarily need physical movement - it can be editing and cropping instead?

JK Liu: Right. One is ego motion; the other is focus movement, cropping within the panoramic image you've already captured.

On the whole, moving a drone is still much simpler than moving something on the ground, because above a certain altitude there are far fewer obstacles.

Jack Zhang: So a bipedal Insta360 device walking around the house shooting isn't the right direction to imagine?

JK Liu: In a complex scene like that, traversability is quite a challenge. A robot vacuum already has to solve a lot of problems; an indoor shooting robot has to solve far more than a vacuum does.

Jack Zhang: There's a huge amount of capital and a lot of talented people exploring embodiment right now. For their data needs, the devices of the past aren't necessarily the optimal way to collect anymore. Are there organizations like that talking to you? How do you see it?

JK Liu: Broadly, in this space we think the path for data collection hasn't converged yet. Some collect first-person data, some collect first-person plus third-person, some collect panoramic, and some don't collect data at all - they learn straight from video.

To put it bluntly, it hasn't reached scale yet. There are only a few points of consensus right now - for instance, that you don't use teleoperation in the pretraining stage, because the collection cost and the data volume can't keep up.

Jack Zhang: Could that change the nature of demand? Everything imaging-related we've built so far has been for people. If the age of robots arrives, the demand for sensory data for robots could be even larger than for humans - the way human search volume is falling today while AI search is rising exponentially. If the embodied era arrives, how do you think imaging changes?

JK Liu: I think the impact on our industry's target customers is fairly small, whether from AI or from embodiment. Because imaging serves people's record of their lives, and recording has nothing to do with AI. It's about the beautiful, precious moments of a person's own experience.

The things we shoot on our phones aren't all meant to be shared. Some are just for our own memories, or to put in a digital photo frame at home. People don't tend to generate an image of something that never happened. From the standpoint of recording, I don't think there's much impact.

But beyond recording, imaging has two other kinds of demand.

One is sharing. Here AI can act as an amplifier, bringing out effects you couldn't have captured yourself. The other is creation - using imaging to express an idea, or to shoot a commercial or make a piece of work. There, I believe AI can also play an important role.


04 · The Cameraman needs a dedicated brain - one with taste

Jack Zhang: There's a debate in the industry lately about whether an all-in-one device will eventually "rule them all," and whether a product built for one narrow niche will always have a low ceiling. It rhymes with a question in AI - whether to build a general-purpose agent, or to go all the way on a few well-defined needs. How do you see it?

JK Liu: I think the camera category is fairly special. From first principles, imaging content falls into a few different types, and that alone guarantees the hardware will split into at least several forms.

First, shooting what we ourselves see. For that we'll mostly use a phone, a camera, or a gimbal camera.

Second, POV - first-person view. This kind of content is common during motion, and it demands a wearable camera, so the form becomes glasses, a thumb-sized camera, or an action camera mounted on your head.

Third, the third-person view, which frees your hands. From the Cameraman's standpoint it definitely isn't something you operate; it's independent. So beyond what you hold in your hand, there's what you wear on your body, and there's the thing at some distance from you that moves autonomously and shoots you.

Jack Zhang: So a single all-in-one device won't necessarily realize the whole Cameraman?

JK Liu: There's an important concept in imaging called focal length - how wide and how far a frame can capture. Different focal lengths serve different purposes.

Panoramic is great for shooting everything around you from your own vantage point; you capture it all and let computation and AI handle it afterward.

But from the standpoint of recording, a good portion of your frames has to be about shooting you: the camera up close on your face with the background, farther out on your full body with the background, or farther still with a telephoto on your upper body with the background - each produces a different feeling. And there's a category of beautiful video shots whose common trait is the relative motion between subject and background, which only a third-person view can produce.

So imaging has two broad forms and needs: one is shooting you from another vantage point, the other is shooting other things from your vantage point. Both are important pieces of visual language. For data collection you need both; for applications they're two different kinds of demand. A camera may be like a car: is its essential function getting from point A to point B? Yes - but it still branches into a great many things.

Jack Zhang: So from the brain's perspective - in embodiment everyone's saying you need one general-purpose brain. For imaging, is the future brain one, or many?

JK Liu: Just for the Cameraman's brain, our current read is that it probably won't be an especially complex thing. Could we be absorbed by a larger model that swallows this piece whole? That's possible too.

But there's another line of reasoning from the business side. I believe that whether you're building models or building intelligence, commercially everyone will still pursue their own distinctiveness - the part that only you have.

So by default, it's genuinely possible that one general-purpose brain does everything across every industry well, or that for photography and videography specifically, one brain handles most scenarios quite well.

But from the standpoint of human nature and of business, I think it's very likely everyone will want this: I understand this scenario deeply, so I've trained on more data here, some of it data only I have, and maybe the shooting method itself is something I invented - I turn that into a model or a skill dedicated to this scenario, and I charge a premium for it.

Jack Zhang: When it comes to imaging, what may matter more to us about that brain is better taste - in theory it should surpass mine, if it can deliver that. With what data and what methods can you make sure it reliably delivers higher taste across the board? Is there a clear answer to that today?

JK Liu: That's a great question. The direction we're training on for composition and editing is still common-denominator taste.

The reason our auto-editing - the Cameraman - is relatively deliverable today is, first, that it doesn't act back on the physical world, and second, that it can offer several options at once. Its evaluation standard doesn't converge, so one of them usually lands. The problem is you can't guess the customer's intent 100% of the time.

Internally we have a formula: customer satisfaction equals what you deliver divided by their expectation. But the expectation itself is hard to pin down. As an analogy, say an AI decides what you eat tonight. However well it knows you - the way your wife decides what you're having for dinner - it can't nail exactly what you're in the mood for every single time.

The next direction probably comes back to the customer: we offer several options, learn their preferences from which ones they pick, and feed that back as context or parameters for the next inference, until we get to something personalized for each individual. But that hasn't played out yet, because how often a customer edits videos is nowhere near the volume of data from how often they scroll through them (laughs).

So on taste - on this kind of demand - the hit rate for matching customer needs always seems to have an invisible ceiling.

Jack Zhang: I know you already have a feature that edits videos for users directly in the cloud, and some users are even paying for it. Do you think that share of the service will keep growing as the technology and the taste improve?

JK Liu: I think it's inevitable. You've already spent all that time shooting, so of course you'll want it summed up, whether for yourself or to share with friends - and it takes a real amount of time. You've already spent the money to travel and the time to shoot; from the standpoint of the final step, the demand is clearly there. In the past, editing was simply too costly for customers, so many people skipped that last step.

We currently have some premium editing services, at around 6 RMB per clip. In the roughly two months since launch we've edited hundreds of thousands of clips, and that's still just the limited rollout.

Jack Zhang: How many rolls does it take before they keep one?

JK Liu: Our export rate is above 50% right now. For every two options we give, roughly one gets picked - about that level.

Jack Zhang: One in two gets used.

JK Liu: Yes. But that still depends on where we sit. We judge this to be a must-have, and in the end some people will probably choose to buy the compute outright - after all, at 6 RMB a clip with some chance of having to roll for it, a fair share of people just don't consume that way. It may eventually have nothing to do with buying compute either: as phone compute grows, it might run entirely on the phone, or on a NAS at home.

Jack Zhang: Do you think where you "sit" could change in the future?

JK Liu: Overall, customer behavior and consumption habits are a distribution; it will never be 100% one thing. Not 100% cloud, not 100% buying the compute on the camera, not 100% on the phone or a NAS.

From a company-strategy view, we will consider extending computation from one platform to others. But a big part of it is the maturity and penetration of the platform itself. Charging a customer is a serious matter today; we have to make sure the experience is good. And what can guarantee a consistent experience today is either the cloud or your own hardware. Third parties - whether a NAS or a phone - are still far from our compute target. There's basically no consumer-grade chip or product delivering hundreds of TOPS integrated into a phone or a NAS. But once that trend takes off, our compute service will certainly extend along with it.

Jack Zhang: The impression that stuck with me from this whole conversation is that if we're very confident about AI's intelligence today, we probably shouldn't be over-confident about taste, because taste seems to be personal. So at the level of the Cameraman, if taste is the key, then it's hard.

JK Liu: Let me tell you a joke: every carmaker has great designers, but whether the car looks good really comes down to the taste of the person running the company.

Jack Zhang: In the end it still comes down to the boss. Thank you, JK. Here's hoping your taste stays on point and brings us more great products.

If you read both feeds, follow @GeekParkHQ.

Get the China angle before it hits the global feed.

@GeekParkHQ

Subscribe by email

Get a note when GeekPark publishes a new English story.