The science fiction writer William Gibson once said, “The future is already here – it’s just not very evenly distributed.” Companies like Google, Amazon, and Facebook are investing billions to advance deep learning technologies and bringing science fiction to their billions of users. They are also giving away most of the infrastructure and algorithms that they create for free, allowing businesses across the world to get in on the AI action. The availability of world-class AI research is arriving at an incredible pace and can be deployed for pennies. As this trend accelerates, what does that mean for the rest of us as both users and consumer product creators?
From “Can We?” to “Should We?”: Emerging Challenges in Consumer Software































Auto-generated transcript - may contain errors. Tap a timestamp to jump the video.
So a picture is worth a thousand words. It's an expression we've all heard probably a thousand times. And I just wanted to take for a minute to get everybody to to get their head around that and and start to think about something. I'm gonna share share something quite quickly here.
So first, if you see this word. Or you might look at it. See most people can read it. You understand it. You might think about something. You might think that's what happens when the sun goes down. But there's not much beyond that. And then if you compare it and look at something like this.
And you think about even in the same amount of time, it will conjure different feelings. You start to think about what's going on. You start to think about where do I get to to go enjoy a sunset like that relative to the beautiful summer weather we're having here in Edinburgh today.
You think about the last time I was on a beach, the next time I'm going to be on a beach, the next time I'm going to get to a holiday, the next time I'm going get to relax. All of that gets conveyed in just the same amount of time, far more than what you can get in a in words.
You know, we read it about three hundred words a minute, so a thousand words is about three and a half minutes versus three seconds. So I really just want to talk a little bit about the power of images and and what does this mean as this evolves from a from an artificial intelligence and machine learning perspective.
And really provoke a couple different ideas and thoughts about what we can do and and how we need to think about that. So just a couple more images. If you just think about the density of information that we consume when we see these.
We look at this, we might recognize the sport, might recognize the event. This is Muhammad Ali Nakhanjo Frazier. You might recognize who the person is. You can imagine this might take place in a stadium. You can see the the emotion and the relief and the victory and intensity that happens just by looking at a photo.
And imagine trying to write a news article about this. You would have to spend far more time for somebody to consume that than what they can get in a picture. Well, look at something like this. It's a relatively famous picture of Albert Einstein.
But you might see something, it looks fun, you might recognize who it is, you might start to think about what does Einstein represent in terms of science. And the amount of information that we can not just process and synthesize, but what we recall.
The connections that we make in our in our brains. Images are incredibly, incredibly dense and we can process them incredibly quickly. So what does that mean when we start thinking about, well, what happens when computers start to see as well or better than we do as humans?
What happens when machines are able to recognize and extract all of that same context and yet be able to look up all of the related context at the same or greater speed. Well, for one thing that happens, cameras start to replace keyboards. Anybody who's ever tried to search for something on your phone, you're trying to search for something really, really specific or find it, you're sitting there tapping, tapping, tapping.
It takes a while. It'd be a lot easier if you could just take a picture of it. So I want to share maybe a couple examples of where where we actually see cameras starting to replace keyboards in in things that we may do every day.
This is the Amazon app. You can go into the app and rather than typing in what you're looking for, you can just take a quick picture. It will recognize it and pull up and let you order that just like that. So you just tap that. You don't even have to type anything.
Interesting thing, this feature has actually been in the app for about three or four years at this point. Not highly celebrated and something that they continue to work on and continue to tune. You look at something like Pinterest. Pinterest released an application earlier this year called Lens.
Pinterest Lens can look at a picture, specifically food in this case, recognize what it is, recognize where you may want to use it, recognize recipes that you can make. These are all different pins. You can click onto these, you actually get various recipes of what you may want to do with with strawberries.
Interestingly enough, by the way, anybody who watches the Silicon Valley TV show, so they had an episode about something called Shazam for Food and then Pinterest released this about three days later. They just happened to be conveniently working on it. Great YouTube clip by the way, you just look up Shazam for Food on Silicon Valley.
So those are kind of neat tricks you can do in apps, do on your phone. But what does this really start to mean when you see somebody taking this to the extreme? When you see people pushing the leading edge of what's actually possible.
And so this is a quote from William Gibson who would, which I just love, which is the future is already here. It is not yet evenly distributed. So I wanted to share a video of what what actually is possible today. This is going to be a video from Amazon.
Some of you may be familiar with this concept, some not. But I thought I'd share this quick video and we can kind of get a sense for for really what what is possible and what is this world that we're beginning to emerge into.
Four years ago, we started to wonder What would shopping look like if you could walk into a store, grab what you want, and just go? What if we could weave the most advanced machine learning, computer vision, and AI into the very fabric of a store so you never have to wait in line.
No lines, no checkouts, no registers. Welcome to Amazon Go. Use the Amazon Go app to enter, Then put away your phone and start shopping. It's really that simple. Take whatever you like. Anything you pick up is automatically added to your virtual cart. If you change your mind about that cupcake, just put it back.
Our technology will update your virtual cart automatically. So how does it work? We used computer vision, deep learning algorithms, and sensor fusion much like you'd find in self driving cars. We call it Just Walk Out technology. Once you've got everything you want, you can just go.
When you leave, our Just Walk Out technology adds up your virtual cart and charges your Amazon account. Your receipt is sent straight to the app, and you can keep going. Amazon Go. No lines. No checkout. No. Seriously. It's pretty cool. Right? You know, when you look at what's actually possible today.
By the way, this store is open in Seattle. It's open to all of Amazon employees right now and will open to the public. So it's only limited to about four hundred thousand people that can use it. Have talked to a number of my friends that that have used it and shop there, it it really does feel like magic the first time that you do it.
And I think that's what's really special about some of the times in in computing that we're walking into is, we're starting to get these things that really astonish us. Things that we may have thought about or read about from a science fiction perspective, just even a few years ago, actually show up and show up in the real world and become available.
So then when you look at something like this and you start to think, these technologies are actually out there. The other thing that's happening is the distribution cycle of these innovations is rapidly accelerating. So if we talk about the new normal for research, you have a number of companies around this world putting in billions and billions of dollars to create all of this leading edge modern science fiction magic in technology.
So here's an article from Microsoft. This got published last October. And they achieved true parity for voice processing to text at the same level of human recognition. Same speed, same level of background noise, same level of accuracy. And there's all kinds of math models behind this.
What what I find is actually really interesting is this thing right here. So if you read that little blurb, they used a piece of software that they had custom developed that is just open source out on GitHub. Anybody in the world can get access to this technology.
And this is the pattern that you see from every single company. You see companies like Google, Facebook, Microsoft, Amazon. Companies in China, where China and and the US are two of the leading locations for AI work. Companies like Baidu, Tencent. You you see all of these innovations coming out there.
And this is really what's happening in research today. The world's leading edge research is available for free and for everyone. And the speed that these innovations will come to the rest of the world and the ubiquity that they'll begin to show up in all of the various applications and services that we use every day is unlike anything that we've seen up till now.
I want to share another example with you. Now, we're going to see this is from a this is from a research team. I forget which university, but here in Europe. And if you look at the bottom left, that's your control case. Staring dead straight.
You look at the one on the on the upper left, somebody just talking and saying some stuff. Now, what's fascinating about this is that through computer vision they're able to extract the mouth motions and project them on to the person on the lower left and that's what you see on the right hand side.
In real time, zero loss, totally believable. Now, what happens when people start to think about what else could you do with this technology if you could take somebody's picture and put somebody else's mouth movements and somebody else's words on it? Well, you get this.
This comes out of a research group at the University of Washington. They put together, it's a great YouTube clip if anyone wants to go check it out, just search for Barack Obama fake speech. It's like a two or three minute speech. And if you close your eyes, it sounds like a Barack Obama speech.
It is the right cadence, the right pitch, the right words. All the vocal inflections, all the pauses, everything is there. But these are the types of things he's saying. You might recognize me, but this is completely fake. And then we start to think about, well, what happens when everybody has access to this technology?
What happens when everybody can figure out what's both optimistic and potentially pessimistic or manipulative things that they can do with it? So it really starts to provoke a question for all of us. Who who can we trust? How do we know if something really becomes real when the availability and ubiquity of technology can show us things that no longer can we believe things that we see with our own eyes?
Another interesting quote from Eric Schmidt from Google. What's pretty interesting to me about this, this is from twenty eleven. Facial recognition was the only technology that he claims Google had built and then decided not to continue. This is also interesting. This is from a company that commercializes some various deep learning and machine computer vision technology.
And you can see there they put together this comparison chart to tell you why their thing is great. I just pointed out the face recognition piece. Because there will be more than zero companies that exploit all of these both for positive and negative.
And there are definitely positives to facial recognition. And this is an article from a couple of weeks, couple months ago, couple weeks ago about the UK making an arrest. You have interesting yet, perhaps questionable techniques. This is about an app called Findface. Exists in Russia, basically works against the Russia Facebook.
But you can take a picture of anybody you see on the street, take a screen cap of anybody you see in a YouTube video, anybody you see online, and it will locate with seventy, eighty, ninety percent accuracy that person's profile on VK, which is the Russian equivalent of Facebook.
Okay. Kind of creepy. You see the headline there questioning what does this really mean for anonymity? You can get to start doing things like, well we've been doing machine learning for forecasting and predictions for a while. We can now do facial recognition and put two and two together.
And you see China now talking about using facial recognition and AI to actually predict what crimes are going to happen before they happen. And in some cases take preventative measures. It's a fair question. Where do you draw the line? Should we draw the line? How do we do that?
So the more time I spent thinking about this, this is kind of how I started to feel. I wonder if we're all just screwed. Right? Over time, will there just be sinister actors that do every bad possible thing with every available piece of technology and no matter what happens, somebody will do something bad with it.
And the more I started to think about it, I don't want to get too philosophical, but I think broadly we have to believe in the positivity of people and of humans and that and most people want to do the right thing. And I'll tell you where I really got some some inspiration around this.
So I was, there was a talk a couple of weeks ago and I'll share a video in a second at a UN forum on what does AI mean going forward? And I start to think about five years, ten years, twenty years. And they talked about a term that really, really stuck with me.
And so the way I kind of translated that back as an engineer is I thought about, well engineers we often look to build things. We often ask, can we do it? Can we put that together? Can we make that work? Can we can we put somebody else's words into Barack Obama's mouth in the right in the right pitch?
But there's a balance here, which is about ethics and ethicists. And ethicists really ask, should we do that? So I just want you to think about that because I think there's some some inspiring and optimistic activities happening because this is not just a conversation happening in engineering.
You actually see this conversation permeating virtually every industry, every vertical. You see business leaders talking about this all around the world. This is the CEO of Audi and I just want to share, this is just a small one minute clip of a talk that he gave for the UN called AI for Good.
Again, there's a YouTube clip on this about fifteen minutes. It's really worth checking out the entire talk. But I thought this particular section was really, really impactful. When we let people try out our research card check, we often see that minute after minute, people gain confidence and trust in piloted driving.
Seeing is believing. However, ethical concerns exist and we take them serious. The best known example of these ethical questions is a dangerous traffic situation where an accident is unavoidable. Imagine a situation where the autonomous car has got three choices: Either it steers left and harms an elderly lady, or it steers right and it hits a pregnant woman, or it drives straight into an obstacle and thus harms the own passenger.
In such a situation, human beings like you and me have no time for thoughtful decisions. We simply react. But interestingly, we expect the autonomous car to make always the right decision. And quite understandably, people are emotionally touched when thinking of such a scenario.
So it's an interesting question. Imagine you were writing software for for any random problem on the web. You might say somebody provides an input value and you want to do the right thing and you write some if statement or or some case statement.
Okay. If left, go left. If right, go right. If somebody types, if somebody sends in banana instead of left or right, you go, oh no, throw an exception. If you're writing code for a car. If you're writing code for facial recognition at immigration, at customs, at visa processing, for the police.
What's the exception that you throw? How do we think about the impact of the code that we write and how does this actually start to impact the real world as these technologies in the real world become more and more fused? So this really got me thinking from a technology perspective, from an engineering perspective, what do we, what should we do?
And I don't know this is an exhaustive list, but I find there's a couple things that guide me in. And something that's really been rattling around in my head for a while is this notion of ethical engineering. And all of us here all have access to influence and create technology products that are going be used by others.
And what are the types of things that we can be doing to ensure that we're doing everything we can to use technology for the better? That we can avoid accidentally creating a negative impact or intentionally creating a negative impact for people. How do we think about user trust and that it takes forever to earn trust?
And just think about your friends. I'm sure everybody has that friend who was a great friend for a number of years. And that one time that they broke your trust, that one time they screwed you, that one time they went behind your back.
You don't forget that. Right? That level of trust is effectively never repaired. And the way that our users use our product, whether they're users or other people at our company, consumers, enterprises, that level of trust is similar. How do you ensure that you never disrupt that?
That you never break that trust? How do you think about making sure people's personal information is sufficiently protected? How do you make sure that when people give you personal information, they know what they're doing as opposed to to taking it? Right? If you think of, you know, the advertising tech on the on the web today is built by a lot of opt out principles rather than opt in.
Is that the right thing? To be honest, I don't know. These are these are questions that rattle around for me of as more and more powerful technology becomes available. What do we do from an engineering perspective? How do we influence this from a technology perspective?
And I come back to those questions of can we and should we? And this is sort of where I came to is that I think from an engineering perspective, we need to evolve from strictly looking at whether we can solve a problem, whether something can be done.
And look at it equally weighted of should we be solving that problem? Is that the right way to do it? Are we doing it in a way that earns trust for from users? That earns that earns credibility, that protects their information, That ensures that we're using technology and as it further and further infuses into our lives.
That we're using it for a benefit and to take things forward. And we're not advocating to create harm. And with that, I'll wrap it up. Thank you very much.