10th Indian Delegation to Dubai, Gitex & Expand North Star – World’s Largest Startup Investor Connect
Artificial Intelligence

Glass supercharges smartphone cameras with AI — minus the hallucinations


Your phone’s camera is as much software as it is hardware, and Glass is hoping to improve both. But while its wild anamorphic lens creeps to market, the company (running on $9.3 million in new money) has released an AI-powered camera upgrade that it says vastly improves image quality — without any weird AI upscaling artifacts.

GlassAI is a purely software approach to improving images, what they call a neural image signal processor (ISP). ISPs are basically what take the raw sensor output — often flat, noisy, and distorted — and turn that into the sharp, colorful images we see.

The ISP is also increasingly complex, as phone makers like Apple and Google like to show, synthesizing multiple exposures, quickly detecting and sharpening faces, adjusting for tiny movements, and so on. And while many include some form of machine learning or AI, they have to be careful: using AI to generate detail can produce hallucinations or artifacts as the system tries to create visual information where none exists. Such “super-resolution” models are useful in their place, but they have to be carefully monitored.

Glass makes both a full camera system based on an unusual lozenge-shaped front element, and an ISP to back it up. And while the former is working towards market presence with some upcoming devices, the latter is, it turns out, a product worth selling in its own right.

“Our restoration networks correct optical aberrations and sensor issues while efficiently removing noise, and outperform traditional Image Signal Processing pipelines at fine texture recovery,” explained CTO and co-founder Tom Bishop in their news release.

Concept animation showing process of going from RAW to Glass-processed image.

The word “recovery” is key, because details are not simply created but extracted from raw imagery. Depending on how your camera stack already works, you may know that certain artifacts or angles or noise patterns can be reliably resolved or even taken advantage of. Learning how to turn these implied details into real ones — or combining details from multiple exposures — is a big part of any computational photography stack. Co-founder and CEO Ziv Attar says their neural ISP is better than any in the industry.

Even Apple, he pointed out, doesn’t have a full neural image stack, only using it in specific circumstances where it’s needed, and their results (in his opinion) aren’t great. He provided an example of Apple’s neural ISP failing to interpret text correctly, with Glass faring much better:

Photo provided by Ziv Attar showing an iPhone 15 Pro Max zoomed to 5x, and the Glass-processed version of the phone’s RAW images.

“I think its fair too assume that if Apple hasn’t managed to get decent results, it is a hard problems to solve,” he said. “It’s less about the actual stack but more about how you train. We have a very unique way of doing it, which was developed for the anamorphic lens systems and is efficient at any camera. Basically, we have training labs that involve robotics systems and optical calibration systems that manage to train a network to characterize the aberration of lenses in a very comprehensive way, and fundamentally reversing any optical distortion.”

As an example, he provided a case study where they had DXO evaluate the camera on a Moto Edge 40, then do so again with GlassAI installed. The Glass-processed images are all clearly improved, sometimes dramatically so.

Image Credits: Glass / DXO

At low light levels the built-in ISP struggles to differentiate fine lines, textures, and facial details in its night mode. Using GlassAI, it’s as sharp as a tack even with half the exposure time.

You can go peep the pixels on a few test photos Glass has available by switching between the raws and the finals.

Companies putting together phones and cameras have to spend a lot of time tuning the ISP so that the sensor, lens, and other bits and pieces all work together properly to make the best image possible. It seems, however, that Glass’s one-size-fits-all process might do a better job in a fraction of the time.

“The time it takes us to train shippable software from the time we put our hands on a new type of device… it varies between few hours to few days. For reference, phone makers spend months tuning for image quality, with huge teams. Our process is fully automated so we can support multiple devices in a few days,” said Attar.

The neural ISP is also end-to-end, meaning in this context that it goes straight from sensor RAW to final image with no extra processes like denoising, sharpening, and so on needed.

Left: RAW, right: Glass-processed.

When I asked, Attar was careful to differentiate their work from super-resolution AI services, which take a finished image and upscale it. These often aren’t “recovering” details so much as inventing them where it seems appropriate, a process that can sometimes produce undesirable results. Though Glass uses AI, it isn’t generative the way many image-related AIs are.

Today marks the product’s availability at large, presumably after a lengthy testing period with partners. If you make an Android phone, it might be good to at least give it a shot.

On the hardware side, the phone with the weird lozenge-shaped anamorphic camera will have to wait until that manufacturer is ready to go public, though.

While Glass develops its tech and trying out customers, it’s also been busy scaring up funding. The company just closed a $9.3 million “extended Seed,” which I put in quotes because the seed round was in 2021. The new funding was led by GV, with Future Ventures, Abstract Ventures, and LDV Capital participating.



Source link

by Team SNFYI

Facebook is testing a new feature that invites some users—mainly in the US and Canada—to let Meta AI access parts of their phone’s camera roll. This opt-in “cloud processing” option uploads recent photos and videos to Meta’s servers so the AI can offer personalized suggestions, such as creating collages, highlight reels, or themed memories like birthdays and graduations. It can also generate AI-based edits or restyles of those images. Meta says this is optional and assures users that the uploaded media won’t be used for advertising. However, to enable this, people must agree to let Meta analyze faces, objects, and metadata like time and location. Currently, the company claims these photos won’t be used to train its AI models—but they haven’t completely ruled that out for the future. Typically, only the last 30 days of photos get uploaded, though special or older images might stay on Meta’s servers longer for specific features. Users have the option to disable the feature anytime, which prompts Meta to delete the stored media after 30 days. Privacy experts are concerned that this expands Meta’s reach into private, unpublished images and could eventually feed future AI training. Unlike Google Photos, which explicitly states that user photos won’t train its AI, Meta hasn’t made that commitment yet. For now, this is still a test run for a limited group of people, but it highlights the tension between AI-powered personalization and the need to protect personal data.

by Team SNFYI

News Update Bymridul     |    March 14, 2024 Meesho, an online shopping platform based in Bengaluru, has announced its largest Employee Stock Ownership Plan (ESOP) buyback pool to date, totaling Rs 200 crore. This buyback initiative extends to both current and former employees, providing wealth creation opportunities for approximately 1,700 individuals. Ashish Kumar Singh, Meesho’s Chief Human Resources Officer, emphasized the company’s commitment to rewarding its teams, stating, “At Meesho, our employees are the driving force behind our success.” Singh further highlighted the company’s dedication to providing opportunities for wealth creation despite prevailing macroeconomic conditions. This marks the fourth wealth generation opportunity at Meesho, with the size of the buyback program increasing each year. In previous years, Meesho conducted buybacks worth over Rs 8.2 crore in February 2020, Rs 41.4 crore in November 2020, and Rs 45.5 crore in October 2021. Meesho’s profitability journey began in July 2023, making it the first horizontal Indian e-commerce company to achieve profitability. Despite turning profitable, Meesho continues to maintain positive cash flow and focuses on enhancing efficiencies across various cost items. The company’s revenue from operations for FY 2022-23 witnessed a remarkable growth of 77% over the previous year, amounting to Rs 5,735 crore. This growth can be attributed to Meesho’s leadership position as the most downloaded shopping app in India in both 2022 and 2023, increased transaction frequency among existing customers, and a diversified category mix. Additionally, Meesho’s focus on improving monetization through value-added seller services contributed to its revenue growth. Meesho also disclosed its audited performance for the first half of FY 2023-24, reporting consolidated revenues from operations of Rs 3,521 crore, marking a 37% year-over-year increase. The company achieved profitability in Q2 FY24, with a significant reduction in losses compared to the previous year. Furthermore, Meesho recorded impressive app download numbers, reaching 145 million downloads in India in 2023 and surpassing 500 million downloads in H1 FY 2023-24. Follow Startup Story Source link

by Team SNFYI

You might’ve heard of Grok, X’s answer to OpenAI’s ChatGPT. It’s a chatbot, and, in that sense, behaves as as you’d expect — answering questions about current events, pop culture and so on. But unlike other chatbots, Grok has “a bit of wit,” as X owner Elon Musk puts it, and “a rebellious streak.” Long story short, Grok is willing to speak to topics that are usually off limits to other chatbots, like polarizing political theories and conspiracies. And it’ll use less-than-polite language while doing so — for example, responding to the question “When is it appropriate to listen to Christmas music?” with “Whenever the hell you want.” But Grok’s ostensible biggest selling point is its ability to access real-time X data — an ability no other chatbots have, thanks to X’s decision to gatekeep that data. Ask it “What’s happening in AI today?” and Grok will piece together a response from very recent headlines, while ChatGPT, by contrast, will provide only vague answers that reflect the limits of its training data (and filters on its web access). Earlier this week, Musk pledged that he would open source Grok, without revealing precisely what that meant. So, you’re probably wondering: How does Grok work? What can it do? And how can I access it? You’ve come to the right place. We’ve put together this handy guide to help explain all things Grok. We’ll keep it up to date as Grok changes and evolves. How does Grok work? Grok is the invention of xAI, Elon Musk’s AI startup — a startup reportedly in the process of raising billions in venture capital. (Developing AI’s expensive.) Underpinning Grok is a generative AI model called Grok-1, developed over the course of months on a cluster of “tens of thousands” of GPUs (according to an xAI blog post). To train it, xAI sourced data both from the web (dated up to Q3 2023) and feedback from human assistants that xAI refers to as “AI tutors.” On popular benchmarks, Grok-1 is about as capable as Meta’s open source Llama 2 chatbot model and surpasses OpenAI’s GPT-3.5, xAI claims. Image Credits: xAI Human-guided feedback, or reinforcement learning from human feedback (RLHF), is the way most AI-powered chatbots are fine-tuned these days. RLHF involves training a generative model, then gathering additional information to train a “reward” model and fine-tuning the generative model with the reward model via reinforcement learning. RLHF is quite good at “teaching” models to follow instructions — but not perfect. Like other models, Grok is prone to hallucinating, sometimes offering misinformation and false timelines when asked about news. And these can be severe — like wrongly claiming that the Israel–Palestine conflict reached a ceasefire when it hadn’t. For questions that stretch beyond its knowledge base, Grok leverages “real-time access” to info on X (and from Tesla, according to Bloomberg). And, similar to ChatGPT, the model has internet browsing capabilities, enabling it to search the web for up-to-date information about topics. Musk has promised improvements with the …