Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
February 11, 2026
·
Raleigh
Converting Apple’s Sharp ML for your devices
Learn to convert Apple's SHARP ML for 2D to 3D scenes into browser and VR applications, with added OpenGL interactivity for virtual hands.
Overview
Apple’s SHARP project uses gaussian splatting to convert regular 2D images into full 3D scenes much faster than previous models. They released the model as pytorch, but i will show how to convert it to run in a browser or in VR on a VisionPro, and use opengl techniques to extend interactivity
Video
Transcript
Generated 3 months ago
Summary
Generating a talk summary...
View full transcript
Ready to, teach you guys about or talk to you guys about, Gaussian Splatting. Has anyone heard of Gaussian Splatting before or or Radiance Fields? Okay. Cool. So there's a couple of people here.
So a couple of weeks ago was kind of big on, you know, AI, social media. Apple released a model to convert, images to these 3 d Gaussian Splats. And Gaussian Splats aren't new. Apple’s not by any means the first company to do that, but their claim to fame was it was a 1 1 model image input, 3 d world output, and it it executed within just a very short amount of time, a couple of seconds. So, they released it as a PyTorch model, and it, you know, it's not necessarily consumer friendly to to use on your different Apple’s.
But, it's relatively easy, you know, in general to take a PyTorch model and convert it to, CoreML. And CoreML, if you're building Apple devices, iOS, iPad, Vision OS, that's kind of the format that works best on your mobile devices. So that's kind of what this project is about. I took, Apple's, Gaussian Spotting model and converted it to run on the Vision OS, and I actually have a Vision Pro, but I forgot to bring it today. But if you come to the next Raleigh tinkerers, I'll have it and I'll let you all try it out.
So it's really cool once you see it in, like, the full kind of immersive, environment. But if you have a PyTorch model, it's it's you know, to convert it to CoreML, you can actually just use the Torch library. It's sort of a built in capability. 1 of the key things that you would need to do is just create a wrapper function that calls the inference mode of PyTorch. So if you've ever played around with PyTorch or TensorFlow, they have an eval mode for inference that sort of optimizes it for evaluations.
And when you do that, you can more or less call the, convert function of PyTorch, but there's a couple of key things you need to keep in mind if you're if you're doing to CoreML. 1 of 1 thing is kind of managing your input resolution. Let's see. I can't really it's really small there. So you so the model CoreML sort of requires kind of fixed inputs, so you would need to kind of scale your your inputs to match the core ML format.
The other thing that's really important is naming your inputs and outputs. So by default, if you just call it convert, it doesn't do this operation and and it'll fail, or won't fail. But if you try to use it, the the output layer names are gonna be just gibberish. You're not gonna know what's what. So you can give your output names like this 1 has, colors, opacities, quaternions, singular values, mean vectors, which are all kind of things that define what a Gaussian splat, is.
And then, of course, you can name your inputs, which in this case is image and disparity factor. That's really important. And when you con when you call the convert function for PyTorch to CoreML, By default, it's going to want to use the metal library for converting, but that's not actually going to work when you're trying to convert from PyTorch to CoreML. It just gets stuck in the cycle, cycle, and I let it run for probably hours before I realized something was broken. So you wanna set the compute unit to CPU only, and it shouldn't take more than about this is a 2 gigabyte model that Apple released, and it took about 10 minutes to convert it to CoreML.
And then, It took about 10 minutes to convert it to CoreML. And then iOS 17, that's the latest target right now for CoreML outputs. And when you do that, it'll generate you a, CoreML model. Let me see if I can get this. So originally I was gonna show the Vision OS, but I'm gonna try to I created a local demo as well.
I'm gonna try to show that real quick. And then take a look at Vision OS. Okay, it's not opening, but you can, this is the Gausses Spotting. 1 of the unique things about them is they actually are really performant to run and you can actually run them in your web browser with WebGL and it's you get great frame rates. So this is 1 that I converted, just a minute ago, and this is the VisionOS simulator.
So this is kind of what you would see in the VisionOS, the headset. So if you if you look at this 1, it's, from from my back deck. And 1 kind of interesting thing about Splats is if you go to about the perspective of the where the camera was taken, it looks exactly like the photo. It's a pixel sort of perfect documentation. And that's kind of the perspective there.
But because it's taking this 2 d image and converting it to 3 d, you can basically kind of go into, the picture itself and you can see it sort of degrades the more you kind of go into the 3 d world. It kind of degrades a little bit, but that kind of helps you kind of see what the splat is. So once you kind of go step out of the 2 d perspective and you kind of go really close to 1 of these things here, you can kind of see this is essentially what a splat is. It's an ellipsoid with like a color and a shape parameter and a 3 d position. And when you put them all together, this has 1,200,000 splats that that system is rendering.
It looks like a photorealistic picture. So, this process is actually to convert an image to this 3 d representation. It's about a 10 second process on VisionOS, and it works basically on any picture that you can take. There's other splatting libraries that work on what they call 4 d splats, which is like a video system where this representation will actually move. I think that's going to be a bigger thing moving forward.
And then I've got a with the Vision OS, it actually has, interactions to where you can actually use your hands and and bump into the splats. I've got a video of that on the website. It's linked on the tinkerers if you want to check that out. But it's I think what I think the first splat renderer that I've seen that actually responds to your physical touch. The other thing was Let's see.
I think I have 1 other thing to talk about with splats. This takes about 10 seconds to convert the image to 3 d. Yeah. Is it running locally? Yeah.
Locally, the CPU, the laptop, the Vision OS. You can actually run it in the browser because you can convert the model to Onyx as well. But the browser right now, both Chrome and Safari, they hit the model as a little bit too large. The 1.2 gigabyte model is too big for the browser. Yeah.
For this model, it only takes 1 picture input, but there are other researchers who have released ones that do multiple viewpoint model to do oh, that was the other thing. Because of the way the kind of the 3 d system works, when you do that multiple viewpoint model, it's really good for things like, you know, potentially robotics. It could be a good input for, like, your world models and your action models because it can take, you know, it can represent things volumetrically. Robots can see that this is actually a table that it needs to go around and things like that. So I think you'll see that use case for splatting become a little bit more common because it is a very powerful and fast way to represent 3 d world data with photorealistic accuracy.
So, yeah, hopefully next time people come, we'll have the Vision OS. And it's a lot more kind of cool to see when it's immersive in a 1 to 1 environment and you can kind of walk around in it. And so it's pretty neat.
Links
Tech stack
Finding related talks...
Compose Email
Sending...
Email preview
Loading recent emails...