<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Alex Kolchinski</title>
    <link>https://alexkolchinski.com/</link>
    <atom:link href="https://alexkolchinski.com/feed.xml" rel="self" type="application/rss+xml" />
    <description>Writing on AI, startups, and where things are headed.</description>
    <language>en</language>
    <lastBuildDate>Fri, 28 Mar 2025 16:05:05 GMT</lastBuildDate>
    <item>
      <title>Announcing TalkToSales, which makes websites voice-native</title>
      <link>https://alexkolchinski.com/2025/03/28/announcing-talktosales-which-makes-websites-voice-native/</link>
      <guid isPermaLink="true">https://alexkolchinski.com/2025/03/28/announcing-talktosales-which-makes-websites-voice-native/</guid>
      <pubDate>Fri, 28 Mar 2025 16:05:05 GMT</pubDate>
      <description>Today, I’m proud to announce the launch of TalkToSales (www.talktosales.com), which pioneers a new way people can interact with computers. When you watch science fiction movies like Star Trek, people are often shown talking to computers rather than typing to them. And when visuals like maps and diagrams are useful for the conversation, the sci-fi […]</description>
      <content:encoded><![CDATA[<p class="wp-block-paragraph">Today, I’m proud to announce the launch of TalkToSales (<a href="http://talktosales.com">www.talktosales.com</a>), which pioneers a new way people can interact with computers.</p>



<p class="wp-block-paragraph">When you watch science fiction movies like Star Trek, people are often shown talking to computers rather than typing to them. And when visuals like maps and diagrams are useful for the conversation, the sci-fi computers show them exactly when they’re relevant.</p>



<p class="wp-block-paragraph">This makes a lot more sense than the way we currently interact with software.</p>



<p class="wp-block-paragraph">As humans, we’re wired to exchange information by talking and listening, not by typing and reading. But our interactions with computers to date have largely been limited to typing and reading, because until last year, computers weren’t able to have spoken conversations with us well enough that it felt natural. So, we’ve had to build software that relied on interactions with mice, keyboards, and touchscreens, instead of more natural spoken conversations.</p>



<p class="wp-block-paragraph">These constraints were at their most extreme a few decades ago, when underpowered computers and displays forced us to interact with software through command-line interfaces (CLIs). But as computers and screens got better, most software quickly moved to easier-to-use graphical user interfaces (GUIs). GUIs are a lot better than CLIs, but they can still be awkward — navigating nested menus with hundreds of buttons is hard and slow. This is especially the case on phones, where small screens make it even harder to interact with complex software.</p>



<p class="wp-block-paragraph">But now, just like the graphical part of the Star Trek computer was unlocked a few decades ago, making navigation through GUIs possible, the voice part was unlocked last year, with AI finally able to have natural spoken conversations with people. And now that the “Star Trek computer” is finally possible, we have the opportunity to reimagine how software should work — both how we can use it by talking to it, and how the graphical parts of it should look now that they can be focused on displaying information instead of showing buttons and other things for users to click on and type into.</p>



<p class="wp-block-paragraph">But so far, we’ve barely begun exploring what has become possible with voice-controlled software.</p>



<p class="wp-block-paragraph">The AI apps themselves — primarily ChatGPT — have allowed people to experience talking naturally to AI. And a number of companies have begun automating business phone calls with AI systems as well. But very little has been done so far in integrating rich, graphical software with voice AI in ways that realize the vision of the “Star Trek computer”.</p>



<p class="wp-block-paragraph">We believe one major place this needs to happen is the Web.</p>



<p class="wp-block-paragraph">Of course, desktop software still exists and would be far easier to use with voice controls, and smartphones would be much nicer to use with smart voice control as well. (Hurry up, Siri 2.0!) But ever since the Internet and cloud software revolutions happened, much of humanity’s interactions with software — and our interactions with each other, when mediated by software — happen through the Web.</p>



<p class="wp-block-paragraph">So, the question we set out to answer with TalkToSales has been — “How should the Web work now that it’s possible to talk to computers?” This has led us down some paths that were quickly obvious, and down others that we only uncovered along the way.</p>



<p class="wp-block-paragraph">Initially, like many entrepreneurs today, we were focused on automating business phone calls. This is a niche where voice AI can create value very quickly by replacing “press 2 for customer support, press 3 for returns” type systems and handling repetitive calls that human call center workers would otherwise have to spend time on. </p>



<p class="wp-block-paragraph">But we realized that in many cases, these phone calls are only happening in the first place because people aren’t able to get their needs met through a website. This could look like anything from an e-commerce customer with a question about a product needing to dial a support number, to a businessperson shopping for software needing to schedule a sales call to learn enough to evaluate a product. In any case, this kind of “wait on hold/schedule a call” flow is awkward, time-consuming, and frustrating.</p>



<p class="wp-block-paragraph">And in the case of a website telling visitors to call a phone number to accomplish something, it’s also bizarre to see a 1990s technology (websites) have to redirect users to an 1800s technology (telephones) to get something done.</p>



<p class="wp-block-paragraph">All too often, the hassle of getting on a call with a person &#8211; and the knowledge that it’ll be impolite to leave early even if the call is a waste of time &#8211; deters people from getting on a call at all, leading to unfulfilled needs for them and lost sales for businesses.</p>



<p class="wp-block-paragraph">We realized something here wasn’t right — why were we working on using AI to automate the phone calls that people were often placing when a website couldn’t meet their needs, when we could work on making the phone call unnecessary in the first place?</p>



<p class="wp-block-paragraph">Why shouldn’t we instead turn the website INTO the call?</p>



<p class="wp-block-paragraph">Thus, TalkToSales (T2S) was born.</p>



<p class="wp-block-paragraph">What we are launching today is the first product that can make a website have a real conversation with visitors. The conversation can use any AI voice you want, and even a photorealistic video avatar (our demo uses my face and voice), an animated character, or no video representation at all if you want to keep it simple. T2S can also be used in text input and/or output mode for users who are in quiet places or want privacy.</p>



<p class="wp-block-paragraph">T2S works with desktop browsers as well as with mobile ones — where it can be particularly useful in helping users have a rich interaction with a website despite the limitation of a small screen. </p>



<p class="wp-block-paragraph">T2S bolts on seamlessly to existing web pages and can give users live tours of a page during a conversation, scrolling and navigating around to show content as it becomes relevant. It can even navigate around complex web apps while talking to users, allowing for live customer support, onboarding, etc. — or other use cases like helping users comparison shop on an e-commerce page.</p>



<p class="wp-block-paragraph">As well as navigating around websites, T2S can show off dynamic slides, including visual assets like images, videos or even things like animated 3D diagrams, as they become relevant to the conversation. The materials available as visual assets can be customized to each use case.</p>



<p class="wp-block-paragraph">T2S is also configurable with custom knowledge bases, style guides, and more, so that the AI can execute faithfully on any company’s goals and stay on-brand. This includes the degree of “improvisation” allowed, so that companies in sensitive industries can ensure the AI doesn’t make anything up, but less sensitive use cases can allow more freedom for a wider-ranging conversation.</p>



<p class="wp-block-paragraph">T2S also includes the ability to book conversions right inside a conversation. For a software product, a conversion might be a prospective customer booking a follow-up call with a salesperson. For e-commerce, it might look like an immediate purchase. In any of these cases, a customer can commit to a call, purchase, or other conversion event without even leaving a T2S conversation. Time kills deals, and T2S ensures that appropriate conversions are offered proactively to customers before they have a chance to leave the page and forget to follow through.</p>



<p class="wp-block-paragraph">T2S also supports qualifying leads right through the conversations it has with visitors. It would be counterproductive to encourage every single visitor to book a sales call, and waste a sales team’s time with low-intent or poor-fit leads. It would also be counterproductive to talk to a prospective investor as though they were a prospective customer, or to talk to a prospective recruit as though they were a journalist.</p>



<p class="wp-block-paragraph">T2S supports asking visitors relevant questions to learn what kind of interaction it makes sense to have with them, and driving them to the conversion event that’s most appropriate for them, whether it’s an immediate purchase, a follow-up call, a mailing-list signup, or nothing at all. Websites are one-size-fits-all, but a T2S interaction &#8211; including the dynamic content as well as the conversation itself &#8211; can be completely customized to each visitor.</p>



<p class="wp-block-paragraph">Moreover, the dynamic, flexible nature of a T2S conversation doesn’t just make for a better experience for web visitors and more conversions — it can also be hugely beneficial from an analytics perspective. Traditional web analytics are extremely limited by the constrained nature of a user’s interactions with a traditional website. When a user uses a GUI web app, the main trace they leave is the pages they visit, the time they spend on them, and where and when they scroll and click. That can only tell you so much about what people are actually hoping to get from your website, and if they’re getting it or not.</p>



<p class="wp-block-paragraph">The advantage of T2S is that having a real conversation with someone is inherently much more informative than recording how they read a “brochure” style web page. And T2S offers full transcripts of users’ conversations &#8211; updated live as they happen &#8211; so it’s possible to see exactly how visitors are interacting with a page in real-time. By seeing what visitors ask about, seeing what answers and content are offered to them in exchange, seeing if they express satisfaction or frustration, and why, it’s possible to get a deep, rich view into customer preferences and needs, including which needs are being unmet.</p>



<p class="wp-block-paragraph">T2S AI is also able to automatically identify trends in users’ conversations and surface problems, unanswered questions, and unmet needs, offering a constant finger on the pulse of a business. In this way, T2S essentially becomes an always-on user research interview with every single user who opts into the conversational experience on a business’s website, yielding insights for the business’s leaders that are comprehensive, nuanced, and always up-to-date.</p>



<p class="wp-block-paragraph">In addition, T2S offers one other feature that we believe will change how people interact with each other over the Internet: when a web visitor is having a conversation with T2S, a representative from the company behind the web site reading the live transcript can choose to “call” the visitor. If the visitor accepts (in voice-only mode, or with video on), the company’s representative replaces the AI avatar and is able to have an audio or video call with the visitor instantly. In this way, the company’s founders, salespeople, etc. can actually jump in live and interact with users as they browse the website. </p>



<p class="wp-block-paragraph">This is exactly what physical-world shopkeepers have always done, but it has never before been possible with online businesses.</p>



<p class="wp-block-paragraph">Now, customers will be able to make an instant connection with the people behind a T2S-powered website, which we expect will hugely increase the number of customer conversations businesses can have. For businesses for whom the rate at which web visits convert to conversations is an important factor in the performance of their sales funnel, we expect this functionality will significantly increase the number of web visitors that make it to a conversation, then all the way down the funnel to a purchase.</p>



<p class="wp-block-paragraph">Among the features of TalkToSales, we see the AI-powered conversation functionality and the ability to transition to a human conversation as being complementary to each other. Web visitors are often deterred from attempting to talk to a real person because getting through to someone is often hard, and even when it isn’t, it’s a big mental jump to go from passively browsing a web page to interacting with a real live person who you have to maintain a certain degree of composure in front of. It’s also intimidating to get on a call with a person that you don’t know you’ll want to stay on, since it’s impolite to hang up suddenly. </p>



<p class="wp-block-paragraph">Talking to an AI is much easier to initiate, and has much lower social barriers, since there are lower stakes to the interaction and it’s not impolite to hang up suddenly. But if an AI conversation goes well, it can be much more natural to transition to talking about the same thing with a human, compared to going straight from browsing the web to a human conversation. In this way, T2S can create a smooth on-ramp from web browsing to the human conversations that drive closed deals.</p>



<p class="wp-block-paragraph">In fact, the benefits go both ways: web visitors can get a better sense of a business’s offerings by talking to the AI before committing to talking to a person, and the company’s representatives can use the live transcripts of AI conversations as a filter to decide which web visitors are worth allocating a human to talk to. In this way, we see T2S as not just an AI guide, but also as a way to let people connect with each other to do business over the Internet more smoothly and efficiently.</p>



<p class="wp-block-paragraph">We think TalkToSales can be useful for a number of use cases out of the box.</p>



<p class="wp-block-paragraph">The one we’ve focused on the most to date — and which you can see in our demo — is for websites of businesses that are selling something, where the website is a significant touchpoint in customers’ purchasing decisions. This could be a software business, where T2S could help users understand the software product and give them a live tour of it — just like what you see in our own demo. Or it could be an e-commerce site, where T2S could help users comparison shop — imagine a Home Depot man in an orange apron helping you shop on their website for specific tools or parts. Or it could be an insurance webpage, where T2S could help users compare policies and get to a purchase — maybe the avatar could be Flo for Progressive or the gecko for Geico. It could even be something else entirely, like an apartment building’s website, where users could take virtual tours of different apartments and ask questions, while being “shown around” by a T2S avatar.</p>



<p class="wp-block-paragraph">T2S could also be useful for post-purchase situations — one we’ve talked a lot about is using it as a guide for complex software products. Imagine Microsoft Clippy (remember that?), but that actually works, and can show users how to use Excel, or Salesforce, etc., and even take actions for the user based on spoken requests — &#8220;Can you please make this row bold?&#8221;, etc.</p>



<p class="wp-block-paragraph">The sky is the limit!</p>



<p class="wp-block-paragraph">Today, you can try a fairly simple example for yourself on our website &#8211; <a href="http://talktosales.com">www.talktosales.com</a></p>



<p class="wp-block-paragraph">Our demo uses a video and audio clone of me, Alex. It qualifies you as a lead for Talk to Sales, and directs you to an appropriate conversion event depending on how it qualifies you. It shows you around our simple landing page as product details become relevant to the conversation, and it shows dynamic slides with images and video sourced both from our library and from Internet stock photos, as relevant to the conversation. For real use cases, we’d expect to be using company-approved libraries of assets rather than stock photos, of course, but bear with us for the demo :).</p>



<p class="wp-block-paragraph">We’ve been working on this demo for a few months now, partially as a way to prove to ourselves that our vision was really possible with today’s technology. It was very hard to pull off with today’s tech, but it’s working, and it’s quite a compelling experience if you ask me — see for yourself! What else is encouraging is that all of the underlying AI technology is improving at a blistering rate, so what you see today is the worst T2S is ever going to be.</p>



<p class="wp-block-paragraph">Now that we’ve proven that this approach works, we’re moving on to commercializing it. Specifically, we’re now looking for three VIP early customers with whom we can deploy TalkToSales on their landing pages (or in a specific flow on their site, e.g. a checkout experience). Our goal is to choose three companies for whom there’s high potential of increasing revenues with a richer, more personalized web experience, and a smoother on-ramp from web visit to human conversation and/or purchase.</p>



<p class="wp-block-paragraph">Since T2S is a brand-new product, we have no data on revenue lift yet — so we’re extremely motivated to make our first three customers insanely successful by lifting revenues by as much as possible.</p>



<p class="wp-block-paragraph">We’re expecting these first implementations to be very hands-on — we’re willing to customize T2S to your specific use case to a significant degree, and we’re happy to handle all the technical dirty work ourselves. We’ll also make our initial implementations low-risk from an economic point of view (let’s talk details if you’re interested) and from a technical one as well — for example, we can show the T2S module to just a small percentage of your web visitors initially, and ramp up from there. And if T2S ever goes down, visitors will still be able to use your website exactly how they could before.</p>



<p class="wp-block-paragraph">If you’re interested in exploring being one of our first VIP customers, please reach out at <a href="mailto:alex@talktosales.com">alex@talktosales.com</a> or book a call at <a href="http://book.kolch.in">http://book.kolch.in</a></p>



<p class="wp-block-paragraph">And if you or your company aren’t the right fit, but you know someone who might be interested or even just curious, we’d really appreciate it if you show them this post and/or our demo!</p>



<p class="wp-block-paragraph">We believe that augmenting websites with natural, conversational experiences that include smooth on-ramps to reaching a human or making a purchase will measurably lift revenues for many companies that do business online.</p>



<p class="wp-block-paragraph">We’re looking forward to proving it with our first customers.</p>]]></content:encoded>
    </item>
    <item>
      <title>The “strategic reserve” exposes crypto as the scam it always was</title>
      <link>https://alexkolchinski.com/2025/03/03/the-strategic-reserve-exposes-crypto-as-the-scam-it-always-was/</link>
      <guid isPermaLink="true">https://alexkolchinski.com/2025/03/03/the-strategic-reserve-exposes-crypto-as-the-scam-it-always-was/</guid>
      <pubDate>Mon, 03 Mar 2025 00:06:04 GMT</pubDate>
      <description>Today, President Trump announced that the US Government would begin using taxpayer dollars to systematically buy up a variety of cryptocurrencies. Crypto prices shot up on the news. This is revealing, as crypto boosters have argued for years that cryptocurrency has legitimate economic value as a payment system outside of the government’s purview. Instead, those […]</description>
      <content:encoded><![CDATA[<p class="wp-block-paragraph">Today, President Trump announced that the US Government would begin using taxpayer dollars to systematically buy up a variety of cryptocurrencies. Crypto prices shot up on the news.</p>



<p class="wp-block-paragraph">This is revealing, as crypto boosters have argued for years that cryptocurrency has legitimate economic value as a payment system outside of the government’s purview.</p>



<p class="wp-block-paragraph">Instead, those same crypto boosters are now tapping the White House for money — in US Dollars, coming from US taxpayers.</p>



<p class="wp-block-paragraph">Why?</p>



<p class="wp-block-paragraph">Crypto has been one of the biggest speculative bubbles of all time, maybe the single biggest ever. Millions of retail investors have piled into crypto assets in the hope and expectation that prices will continue to go up. (Notice how much of the chatter around crypto is always around prices, as opposed to non-speculative uses.)</p>



<p class="wp-block-paragraph">However, every bubble bursts once it runs out of gamblers to put new money in, and it may be that the crypto community believes that that time is near for crypto, as they are now turning to the biggest buyer in the world — the US Government — for help.</p>



<p class="wp-block-paragraph">This shows that all the claims that crypto leaders have made for years about crypto&#8217;s value as a currency outside of government control have been self-serving lies all along: the people who have most prominently argued that position are now begging the White House to hand them USD for their crypto.</p>



<p class="wp-block-paragraph">It also reveals how much crypto has turned into a cancer on our entire society.</p>



<p class="wp-block-paragraph">In previous Ponzi schemes, the government has often stepped in to defuse bubbles and protect retail investors from being taken in by scammers.</p>



<p class="wp-block-paragraph">But in this wave, not only has the government not stepped in to stop the scam, it has now been captured by people with a vested interest in keeping it going as long as possible.</p>



<p class="wp-block-paragraph">Our president and a number of members of his inner circles hold large amounts of cryptocurrency and have a vested interested in seeing its value rise — Trump&#8217;s personal memecoin being a particularly notable example. And many other people in the corridors of power in Washington and Silicon Valley are in the same boat. &#8220;It is difficult to get a man to understand something, when his salary depends on his not understanding it&#8221;, and so some of the most prominent people in the country are now prepared to make any argument and implement any policy decision to boost the value of their crypto holdings.</p>



<p class="wp-block-paragraph">How does this end?</p>



<p class="wp-block-paragraph">Once the US taxpayer is tapped out, there’s not going to be any remaining larger pool of demand to keep crypto prices up, and in every previous speculative bubble, once confidence evaporates, prices will fall, probably precipitously. Unfortunately, as millions of people now have significant crypto holdings, and stablecoins have entangled crypto with fiat currency, the damage to the economy may be widespread. </p>



<p class="wp-block-paragraph">The end of the crypto frenzy would, in the end, be a good thing. Cryptocurrency has a few legitimate uses, like helping citizens of repressive regimes avoid currency controls and reducing fees on remittances. But it has also enabled vast evil in the world. Diverting trillions of dollars away from productive investments into gambling is bad enough, but the untraceability of crypto has also enabled terrorist organizations, criminal networks, and rogue states like North Korea to fund themselves far more effectively than ever before. I’ve been hearing from my friends in the finance world that North Korea now generates a significant fraction, if not a majority, of its revenues by running crypto scams on Westerners, and that the scale of scams overall has grown by a factor of 10 since crypto became widely used (why do you think you’re getting so many calls and texts from scammers lately?)</p>



<p class="wp-block-paragraph">I hope that the end of this frenzy of gambling and fraud comes soon. But in the meantime, let’s hope that not too much of our tax money goes to paying the scammers, and that when the collapse comes it doesn’t take down our entire economy with it.</p>



<p class="wp-block-paragraph"></p>



<p class="wp-block-paragraph"><em>Thanks to Alec Bell for helping edit this essay.</em></p>]]></content:encoded>
    </item>
    <item>
      <title>I’m looking for a cofounder</title>
      <link>https://alexkolchinski.com/2024/02/27/im-looking-for-a-cofounder/</link>
      <guid isPermaLink="true">https://alexkolchinski.com/2024/02/27/im-looking-for-a-cofounder/</guid>
      <pubDate>Tue, 27 Feb 2024 20:53:40 GMT</pubDate>
      <description>Summary I’m looking for a cofounder for my next company. I want to work on AI-powered B2B workflow automation software, but I’m not committed to a specific direction yet. I’m currently working on a workflow automation product in the insurance space that just crossed $10K/mo in revenue. In the past, I’ve been the CEO of […]</description>
      <content:encoded><![CDATA[<figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-4-3 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">
<span class="embed-youtube" style="text-align:center; display: block;"><iframe loading="lazy" class="youtube-player" width="640" height="360" src="https://www.youtube.com/embed/FXzkm3JwcFg?version=3&#038;rel=1&#038;showsearch=0&#038;showinfo=1&#038;iv_load_policy=1&#038;fs=1&#038;hl=en&#038;autohide=2&#038;wmode=transparent" allowfullscreen="true" style="border:0;" sandbox="allow-scripts allow-same-origin allow-popups allow-presentation allow-popups-to-escape-sandbox"></iframe></span>
</div></figure>



<ol><li><a class="wp-block-table-of-contents__entry" href="/2024/02/27/im-looking-for-a-cofounder/#summary">Summary</a></li><li><a class="wp-block-table-of-contents__entry" href="/2024/02/27/im-looking-for-a-cofounder/#my-background">My Background</a></li><li><a class="wp-block-table-of-contents__entry" href="/2024/02/27/im-looking-for-a-cofounder/#my-skills">My Skills</a></li><li><a class="wp-block-table-of-contents__entry" href="/2024/02/27/im-looking-for-a-cofounder/#what-i-want-to-work-on">What I want to work on</a></li><li><a class="wp-block-table-of-contents__entry" href="/2024/02/27/im-looking-for-a-cofounder/#who-i-want-to-work-with">Who I want to work with</a></li><li><a class="wp-block-table-of-contents__entry" href="/2024/02/27/im-looking-for-a-cofounder/#other-preferences">Other preferences</a></li></ol>



<h2 class="wp-block-heading" id="summary">Summary</h2>



<p class="wp-block-paragraph">I’m looking for a cofounder for my next company. I want to work on AI-powered B2B workflow automation software, but I’m not committed to a specific direction yet.</p>



<p class="wp-block-paragraph">I’m currently working on a workflow automation product in the insurance space that just crossed $10K/mo in revenue. In the past, I’ve been the CEO of a YC-backed startup, a PhD student at the Stanford AI Lab, an APM at Google, and a software engineer.</p>



<p class="wp-block-paragraph">I’m looking to either stay CEO and join forces with a technical CTO, or to become the CTO to an exceptional CEO. For more details, read on!</p>



<p class="wp-block-paragraph">If you’re interested, or know someone I should talk to, please reach out — I’m at <a href="mailto:alex@kolch.in">alex@kolch.in</a></p>



<h2 class="wp-block-heading" id="my-background">My Background</h2>



<p class="wp-block-paragraph">I grew up programming, “turning pro” when I sold a Flash game in high school. I worked full-time as a software engineer for a year between high school and college, then earned a BS/MS in computer science at UChicago, concentrating on AI.</p>



<p class="wp-block-paragraph">After college, I worked at Google as an APM, then went back to Stanford for a PhD program. Initially, I was planning to focus my research on AI-powered tutoring software, but came to the conclusion that the underlying AI technology wasn’t powerful enough yet to build what I wanted to build. So, I pivoted to doing AI research in natural language processing and generative models, publishing three papers.</p>



<p class="wp-block-paragraph">I then got excited about startups generally and automating food service specifically, and dropped out of Stanford to co-found Mezli, which I led as CEO. We launched a popular autonomous restaurant, but went out of business after our Series A fundraise fell through. I raised $4M from investors including Y Combinator (but failed to raise the additional ~$10M we needed), led a team of ~30, and ran several functions including finance and marketing. Google Mezli for news, reviews, etc., and you can see a video of the tech here &#8211; <a href="https://www.youtube.com/watch?v=DV2I9XwcEZE">https://www.youtube.com/watch?v=DV2I9XwcEZE</a></p>



<p class="wp-block-paragraph">I spent the latter half of 2023 shopping around Mezli’s IP and shutting down the company. While doing that, I used the newfound free time on my hands to build and launch a couple of products solo. One is a B2C utility app (<a href="http://www.readtome-app.com/">www.readtome-app.com</a>); the other is a workflow automation product in the insurance space that just crossed $10K/mo in revenue.</p>



<p class="wp-block-paragraph">I’m now evaluating whether to keep doubling down on this product or to pivot to a different niche that might be a faster place to grow to $1M+ in annual revenue. As I start that discovery process, I’m also on the hunt for a cofounder.</p>



<h2 class="wp-block-heading" id="my-skills">My Skills</h2>



<p class="wp-block-paragraph">While the previous section probably gives you an idea of what I can do, here’s a more specific breakdown of my skills:</p>



<ul class="wp-block-list">
<li><strong>Entrepreneurship:</strong> I’ve launched multiple products as CEO or sole founder, one reaching $20K in monthly revenue and another reaching $10K/mo. Between those experiences, several other exploratory projects, and spending a lot of time in the startup ecosystem, I’ve developed a pretty good sense for what it takes to grow a company from an idea to meaningful revenue — and more importantly, how to quickly discard the many ideas that prove unviable.</li>



<li><strong>Software engineering:</strong> I’ve built a wide variety of software over the last 22 years (time flies…) and I’m very good at picking up new technologies quickly and getting things shipped. This includes the latest wave of AI tech — I’ve leaned heavily on modern generative AI for my last two products. However, I’m very much not a VP Eng who institutes best practices in a large team — I’m more on the “incur tech debt to get to product-market fit” side of the spectrum than the “pay down technical debt to make a product long-term maintainable” side of it.</li>



<li><strong>AI research</strong>: I spent three years of my life largely focused on AI research at Stanford. While I haven’t trained a custom model in a few years, I still have a good understanding of what today’s AI can do, and of what’s likely to be possible soon. I actually don’t think most startups should be in the business of conducting AI research or even of commercializing techniques directly from research papers, but if that changes, I have the relevant background!</li>



<li><strong>Sales, fundraising, and recruiting:</strong> I group these together because in a sense, they’re all manifestations of sales. I’ve run a sales or sales-adjacent process many times, including pitching hundreds of investors to raise $4M for Mezli, recruiting a number of Mezli’s team members, and now selling B2B software in the insurance space.</li>



<li><strong>Marketing and PR</strong>: I’ve run marketing campaigns for Mezli’s robotic restaurant and more recently for the ReadToMe app. The Mezli campaign included local and national media hits and drove ~1M views and ~10K purchases for our brand. I’ve also had success reaching a broad audience with my blog, including multiple front-page posts on Hacker News and ranking #1 on high-volume Google search terms.</li>



<li><strong>Finance</strong>: I have a good understanding of how the numbers work that make businesses tick, and of how the financial markets work as well. This has come from a variety of classes in college and grad school, an internship on Wall Street, and running the financial side of Mezli.&nbsp;</li>
</ul>



<h2 class="wp-block-heading" id="what-i-want-to-work-on">What I want to work on</h2>



<p class="wp-block-paragraph">In short: AI-powered B2B workflow automation software.</p>



<p class="wp-block-paragraph">Like many others in Silicon Valley, I think that the current wave of progress in AI is a generational shift — we’re likely at the beginning of a change as big as the advent of the Internet.</p>



<p class="wp-block-paragraph">Until now, computers have largely been able to transmit and process information only in rigidly-defined ways: humans have had to define exactly how data could be input into programs, stored and processed by them, and given back to users.</p>



<p class="wp-block-paragraph">That’s now changing in a spectacular fashion. Thanks to transformer-style models, many tasks involving unstructured data (natural language, images, video, audio, etc.) that were previously the exclusive domain of humans can now be partially or fully done by software.</p>



<p class="wp-block-paragraph">This includes a broad swathe of repetitive white-collar work that accounts for many trillions of dollars of value created per year. With AI tooling, humans in the affected industries — and that includes most industries — will be made vastly more productive. An analogy might be the transition from paper spreadsheets to electronic spreadsheets a few decades ago, or paper mail to email more recently.</p>



<p class="wp-block-paragraph">That a lot of this automation is about to happen, and that a lot of economic value is about to be created, I haven’t heard anyone seriously contest. The more uncertain question is which companies are going to create that value and capture a share of the value they create.</p>



<p class="wp-block-paragraph">Some important sub-questions:</p>



<ul class="wp-block-list">
<li>How much of the action is going to be captured by incumbents vs. startups?</li>



<li>Which parts, if any, of the new landscape are going to be owned by a fragmented assortment of companies, vs. a small number of huge players?&nbsp;</li>



<li>How will the domains of different companies be carved up&nbsp;— by industries served, type of technology used, both, or neither?&nbsp;</li>



<li>Will some companies playing in the new AI-enabled landscape have significantly better economics than others, which might in turn look more like services businesses? How will that be determined?</li>
</ul>



<p class="wp-block-paragraph">The answers to most of these questions are currently unknowable, so my inclination is to jump in, get to $1M+ with a narrow “wedge” product that’s quick to sell, then grow outwards from there, responding to inevitable shifts in the competitive landscape of the software industry, including advances in the capabilities of AI, as they transpire.&nbsp;</p>



<p class="wp-block-paragraph">I’m inclined to start out by solving a fairly narrow problem in a specific industry to begin with, as opposed to building tools for other companies to use “horizontally” across industries. This is for a few reasons:</p>



<ul class="wp-block-list">
<li>A well-constrained customer profile can make the sales learning curve faster early on, allowing for a faster ramp-up in revenue.</li>



<li>I’m seeing way more founders jumping into the tooling layer right now than into the application layer. This likely implies less competition in at least some vertical-specific product areas.</li>



<li>The landscape of what’s possible with AI, and which tools are needed to do it, is changing at a very fast pace. Many tooling companies whose products are useful and popular today might be completely obsolete tomorrow. A company whose product is solving an application-layer need is more likely to be able to make use of improved AI tooling seamlessly and continue serving its customers instead of being replaced.</li>
</ul>



<p class="wp-block-paragraph">I’m now on the hunt for the right vertical, and the right niche in that vertical, to get started. It’s possible that my current product in the insurance space will be that initial product; it’s also possible that I’ll find something else that I like better.</p>



<p class="wp-block-paragraph">As I start that discovery process, I’m on the hunt for someone to do it with, so that we can shape the product direction of the company together.</p>



<h2 class="wp-block-heading" id="who-i-want-to-work-with">Who I want to work with</h2>



<p class="wp-block-paragraph">I’m in the relatively unusual position of having a background in both the business and product side of entrepreneurship as well as in the technical side, both in software engineering generally and AI specifically.</p>



<p class="wp-block-paragraph">However, rather than continuing as a solo founder, I’m looking for a cofounder for a few reasons. With the right cofounder,</p>



<ul class="wp-block-list">
<li>Starting and running a company together is more fun than doing it alone.</li>



<li>Two heads are better than one.</li>



<li>Many hands make light work.</li>
</ul>



<p class="wp-block-paragraph">As far as who I’m looking to work with, I can see one of two arrangements working well:</p>



<ul class="wp-block-list">
<li>I stay CEO and a CTO joins me who has a history of shipping software quickly. I focus on selling; the CTO focuses on building and I help build when appropriate. In this arrangement, we’d go through the discovery process together before committing to a direction for the company.</li>



<li>I become CTO to an extraordinary CEO. The CEO sells, I build. Because I’ve been the CEO of a startup that had some temporary success, I have a pretty high bar for this one: the CEO would either have to have had a previous exit as a startup CEO, or be a veteran of an industry that they can immediately start making sales in — or both.&nbsp;</li>
</ul>



<p class="wp-block-paragraph">Either way, personal and professional compatibility are key — we should get along famously, and working together should feel like much more than the sum of its parts.</p>



<h2 class="wp-block-heading" id="other-preferences">Other preferences</h2>



<p class="wp-block-paragraph">There are a few other things I should mention about my preferences:</p>



<ul class="wp-block-list">
<li>I feel very strongly about working in-person together, most days of the week, most weeks of the year, in or near San Francisco. I want to stay in-person forever, and I want to stay in the Bay Area indefinitely, with the possible exception of if we end up serving a customer base that’s heavily concentrated in a different city.</li>



<li>Founding a startup requires a level of intensity that’s alien to most people&nbsp;— when done right, it doesn’t leave room in your life for much else that takes proactive effort. You need to be prepared for this. However, some founders grind to the point of negative marginal returns, and that’s not a good idea either. I strive for, and expect from cofounders, a level of intensity that’s far beyond a 9-5 job, but that does (at least periodically) leave room for recovery and perspective.</li>



<li>At this stage of the game I’d be looking to split equity equally, with one extra share to the CEO to break ties. I also prefer a longer vesting schedule than 4 years to align founder incentives for the long-term.</li>
</ul>



<p class="wp-block-paragraph">Interested, or know someone I should talk to? Please reach out — I’m at <a href="mailto:alex@kolch.in">alex@kolch.in</a></p>]]></content:encoded>
    </item>
    <item>
      <title>Announcing ReadToMe</title>
      <link>https://alexkolchinski.com/2024/02/04/announcing-readtome/</link>
      <guid isPermaLink="true">https://alexkolchinski.com/2024/02/04/announcing-readtome/</guid>
      <pubDate>Sun, 04 Feb 2024 23:18:10 GMT</pubDate>
      <description>Today, I’m announcing the public launch of the ReadToMe app, which turns paper books and other printed text into high-quality audio. I originally built the app as a present to my fiancée, who has a reading disability but loves books. Often, she listens to audiobooks while following along in the same book on paper, but […]</description>
      <content:encoded><![CDATA[<figure class="wp-block-image size-full is-resized"><img src="https://alexkolchinski.com/wp-content/uploads/2024/02/ezgif-7-bd9c149c44.gif" alt="" class="wp-image-436" style="object-fit:cover;width:277px;height:600px" /></figure>



<p class="wp-block-paragraph">Today, I&#8217;m announcing the public launch of the <a href="http://www.readtome-app.com">ReadToMe app</a>, which turns paper books and other printed text into high-quality audio.</p>



<p class="wp-block-paragraph">I originally built the app as a present to my fiancée, who has a reading disability but loves books. Often, she listens to audiobooks while following along in the same book on paper, but some books, especially older and less popular ones, are only available in paper form and don&#8217;t come in audiobook or even e-book form.</p>



<p class="wp-block-paragraph">We looked for a way for her to turn paper books into audio, but all of the apps we found didn&#8217;t do a very good job. Many of them were very good at turning e-books and other digital text into high-quality audio, but made many mistakes when scanning paper books. Common issues included inserting page numbers and footnotes into the middle of sentences and getting words wrong or missing them completely. Overall, there ended up being so many mistakes that the books were very hard to listen to — and that was for the better apps.</p>



<p class="wp-block-paragraph">As a Christmas present, I wrote an early version of ReadToMe for my fiancée. When she found it useful, and I had spare time on my hands while shutting down my last company, I built out the app into its current version — which is what you see here.</p>



<p class="wp-block-paragraph">The app lets you scan up to 20 pages at a time and turns them into high-quality audio, with very few mistakes. The app costs $9.99/month for up to 250 pages/month — sorry it can&#8217;t be free; I estimate the $9.99/month will cover the costs of the pretty expensive AI technology it uses on the back end.</p>



<p class="wp-block-paragraph">Known issues that I&#8217;ll be working on fixing if enough people end up using the app include:</p>



<ul class="wp-block-list">
<li>Scans can take a few minutes to come back as audio, especially when scanning multiple pages at once.</li>



<li>Rarely, scans will fail to come back as audio at all and have to be retried.</li>



<li>The AI I&#8217;m using on the back end will sometimes &#8220;correct&#8221; the wording of a book, especially when an author deliberately uses incorrect grammar.</li>
</ul>



<p class="wp-block-paragraph">I expect the app might be useful for a couple of types of user, including:</p>



<ul class="wp-block-list">
<li>People with reading disabilities or just age-related far-sightedness that makes reading hard.</li>



<li>People who want to seamlessly bounce between reading something on paper and listening to a few pages, e.g. while driving.</li>
</ul>



<p class="wp-block-paragraph">I&#8217;m also looking forward to seeing who else might find it useful.</p>



<p class="wp-block-paragraph">If that&#8217;s you, or you know someone who might find ReadToMe useful, please give it a try and/or let them know! And I&#8217;d appreciate any feedback on things that the app does well or poorly — you can reach me at alex@yaksoft.net</p>]]></content:encoded>
    </item>
    <item>
      <title>2023 Reflections</title>
      <link>https://alexkolchinski.com/2023/12/11/2023-reflections/</link>
      <guid isPermaLink="true">https://alexkolchinski.com/2023/12/11/2023-reflections/</guid>
      <pubDate>Mon, 11 Dec 2023 02:25:09 GMT</pubDate>
      <description>This year, I had to shut down my autonomous restaurant startup Mezli — despite a successful launch — after I was unable to raise the money we needed to scale up. This year, I’ve also re-entered the world of AI, which I’ve been away from since leaving Stanford in 2020. I’m now working on a new […]</description>
      <content:encoded><![CDATA[<p class="wp-block-paragraph">This year, I had to shut down my autonomous restaurant startup Mezli —&nbsp;despite a <a href="https://www.youtube.com/watch?v=DV2I9XwcEZE">successful launch</a> — after I was unable to raise the money we needed to scale up. </p>



<p class="wp-block-paragraph">This year, I&#8217;ve also re-entered the world of AI, which I&#8217;ve been away from since leaving Stanford in 2020. I&#8217;m now working on a new workflow automation startup and have gotten very excited about the opportunities created by recent advances in AI.</p>



<p class="wp-block-paragraph">However, the emergence of true artificial intelligence is also shaping up to be the fastest-ever large economic and social change in human history, and tracing how it may reshape business and society is both intellectually interesting and strategically important for entrepreneurs (and just about everyone else, for that matter).</p>



<p class="wp-block-paragraph">While chewing on the lessons I&#8217;ve learned from shutting down a hardware startup and re-entering the world of AI, I&#8217;ve summarized my thoughts into three essays, which I hope can be useful to the startup community and spark an interesting conversation:</p>



<ul class="wp-block-list">
<li style="line-height:2"><a href="/2023/12/11/the-end-of-knowledge-work/">The End of Knowledge Work</a></li>



<li style="line-height:2"><a href="/2023/12/11/the-end-of-the-software-industry/">The End of the Software Industry</a></li>



<li style="line-height:2"><a href="/2023/12/11/founders-beware-hardware/">Founders, Beware Hardware</a></li>
</ul>



<p class="wp-block-paragraph">A summary of my arguments is as follows:</p>



<p class="wp-block-paragraph" style="line-height:1.5">It&#8217;s likely that AI is now only a small leap away from becoming as good as humans at most intellectual tasks. Even if that&#8217;s not the case, its current capabilities are already good enough to do broad swathes of knowledge work that are currently done by humans.</p>



<p class="wp-block-paragraph" style="line-height:1.5">So, we&#8217;re likely entering an era of knowledge work automation that may parallel the last 200 years of physical work automation, but it&#8217;ll likely happen much faster this time. We&#8217;ll likely see far fewer people doing analytical work over the next ~20 years, just as fewer and fewer people did agricultural and other manual work over the past ~200.</p>



<p class="wp-block-paragraph" style="line-height:1.5">This automation of knowledge work is likely to reduce the software industry to a shadow of its former self, as programming will become largely automated.</p>



<p class="wp-block-paragraph" style="line-height:1.5">However, there are many opportunities to build software startups right now, and software presents a uniquely good opportunity for first-time founders to build companies that reward good execution and provide quick feedback. </p>



<p class="wp-block-paragraph" style="line-height:1.5">Hardware companies, on the other hand, are very difficult for first-time founders because their unavoidable capital requirements make them vulnerable to running out of money, and because their slow cycle times make for a slow education in entrepreneurship. </p>



<p class="wp-block-paragraph" style="line-height:1.5">So, I highly encourage today&#8217;s aspiring entrepreneurs to focus on software, especially because we may be entering the last era when the software industry, and software entrepreneurship, exists in its present form. </p>



<p class="wp-block-paragraph">That said, given the rapid changes in AI capabilities, I also think it&#8217;s especially important for today&#8217;s AI founders to keep an eye on the long-term competitive advantages their companies develop, as technology built on top of today&#8217;s AI may become infinitely easier to build on top of tomorrow&#8217;s AI, negating the technical defensibility of much of the work being done by startups today. Other forms of defensibility are likely to be more important in the future.</p>



<p class="wp-block-paragraph">I welcome all comments, especially those that point out things I missed or refute a claim I&#8217;ve made. Hopefully we can all learn from the debate.</p>]]></content:encoded>
    </item>
    <item>
      <title>Founders, Beware Hardware</title>
      <link>https://alexkolchinski.com/2023/12/11/founders-beware-hardware/</link>
      <guid isPermaLink="true">https://alexkolchinski.com/2023/12/11/founders-beware-hardware/</guid>
      <pubDate>Mon, 11 Dec 2023 02:03:06 GMT</pubDate>
      <description>This essay is part of a 3-part series: Hardware startups are sexy. Building flying cars, nuclear reactors, or electric cars is more tangible and in many ways more interesting than writing software. It’s also possible to tackle a wider range of important problems with hardware than with software alone, from global warming to food security. […]</description>
      <content:encoded><![CDATA[<p class="wp-block-paragraph" style="line-height:0.5"><strong>This essay is part of a <a href="/2023/12/11/2023-reflections/">3-part series</a>:</strong></p>



<ul class="wp-block-list">
<li><a href="/2023/12/11/the-end-of-knowledge-work/">The End of Knowledge Work</a></li>



<li><a href="/2023/12/11/the-end-of-the-software-industry/">The End of the Software Industry</a></li>



<li><a href="/2023/12/11/founders-beware-hardware/">Founders, Beware Hardware</a></li>
</ul>



<p class="wp-block-paragraph">Hardware startups are sexy. Building flying cars, nuclear reactors, or electric cars is more tangible and in many ways more interesting than writing software. It’s also possible to tackle a wider range of important problems with hardware than with software alone, from global warming to food security.</p>



<p class="wp-block-paragraph">For these reasons, many founders and would-be founders dream of starting hardware companies, and some of them actually do. Occasionally, that decision works out reasonably well, but most of the time, it ends in tears, especially for founders without previous successful outcomes. This is due to factors inherent to hardware companies, which the startup community is somewhat aware of but perhaps not enough. In this essay, I’m hoping to make those factors clear, with the hopes of saving other founders future pain.</p>



<p class="wp-block-paragraph">This warning comes out of my own story.&nbsp;</p>



<p class="wp-block-paragraph">I’m from a software background — I started programming as a teenager, studied computer science and AI in college and in grad school, and have worked in different parts of the software industry since high school. But in grad school, while studying AI, I became enamored with the idea of automating food service.&nbsp;</p>



<p class="wp-block-paragraph">I realized that automating fast-casual food service (think Chipotle, Sweetgreen, etc.) could bring down the costs of building and operating restaurants by so much that high-quality meals could be sold at half the price they’re sold for today. To realize this vision, I co-founded Mezli with two friends from Stanford, and a year and a half later, we successfully launched a <a href="https://www.youtube.com/watch?v=DV2I9XwcEZE">fully-robotic restaurant</a>. However, despite selling thousands of meals at near-100% uptime and significantly better unit economics than human-powered restaurants, we were unable to raise more money and were forced to shut down.</p>



<p class="wp-block-paragraph">Some of this was due to the downturn in the funding market — it was much harder to raise money in 2022 than it had been even a year or two earlier. But in hindsight, just being a hardware company made us much more fragile than an equivalent software startup would have been.</p>



<p class="wp-block-paragraph">One crucial factor that I underestimated at the beginning of Mezli was the mandatory requirement for raising escalating amounts of capital to keep making progress with a hardware startup. People sometimes talk about hardware startups being “capital intensive”, but this is only part of the picture — plenty of software companies also invest hundreds of millions of dollars into engineering, sales, and marketing before they become profitable. The difference is that most software companies can adjust the amount of capital they invest in product development and growth as conditions change. If investor money becomes unavailable, a software company can usually lay people off and coast on revenues until the economy improves — or never fundraise again.</p>



<p class="wp-block-paragraph">This is simply not possible for most early-stage hardware companies. Hardware products typically require substantial scale to reach profitability, which means huge up-front investments are necessary before turning a profit becomes possible. This means needing to raise numerous rounds of external funding, with each being do-or-die. The immediate implications of this dynamic are bad enough, but there are second-order impacts as well — even if an investor likes a hardware company’s prospects based on its technology and economics, well-founded concern that a single missed fundraise at any point in the future would be enough to kill the company can easily be a significant-enough factor to torpedo the current fundraise as well — a sort of “Keynesian beauty contest” dynamic that sounds academic but is very real.&nbsp;</p>



<p class="wp-block-paragraph">Even a hardware success story like Tesla suffered from this dynamic and would have folded if not for repeated bailouts by Elon Musk. And for every Tesla, there have been countless hardware companies with equally-promising products but without the deep-pocketed backers needed to see them through lean times.</p>



<p class="wp-block-paragraph">A second crucial drawback of hardware startups, which I think is often underappreciated, is the glacial speed of iteration compared to software. A software company can launch and update its product(s) and go-to-market motion near-instantaneously, allowing for very fast iteration. This is a huge advantage to a startup, allowing for fast contact with the market and evolution of its product(s) in the direction of customer pull — essentially the essence of the “Lean Startup” philosophy. Many software companies have found product-market fit and grown large on the back of this approach.&nbsp;</p>



<p class="wp-block-paragraph">But the value of the ability to iterate quickly is not only to the benefit of the startup itself; it’s also to the benefit of its founders personally. When it’s possible to experiment on a daily cadence, to see efforts succeed or fail, and then to try again the next morning, the rate at which it’s possible to become a more-skillful entrepreneur is very high.&nbsp;</p>



<p class="wp-block-paragraph">Hardware companies present a completely different picture than this when it comes to the speed of iteration. Hardware products typically take months or even years to design, manufacture, and ship, so that much more has to be guessed up-front that is only proven right or wrong by the market a long time later. And with only one or several iterations possible before the company proves a success or failure, the founders have limited opportunities to learn from the experience of bringing a product into contact with the real world. This slower learning curve is especially bad for first-time founders, who need to spend years learning lessons they would have learned in weeks with a software product.</p>



<p class="wp-block-paragraph">And one final disadvantage of hardware is that even if a hardware product proves a huge success, the slow cycle times of hardware mean that scaling up that success and enjoying its fruits takes far longer than it would for a typical software company.&nbsp;</p>



<p class="wp-block-paragraph">Of course, there are some exceptions to these rules — hardware companies where mandatory&nbsp; requirements for capital and long cycle times are less of a factor. A common example are companies that use off-the-shelf hardware to deliver a product differentiated by software. This includes companies that use off-the-shelf drones accompanied by computer vision models to inspect bridges and pipelines, companies that use AI models to get robot arms to pack boxes or weld metal, and many such others. By not having to design and build their own hardware, these companies are subject to fewer — but still many — of the issues that dog hardware startups.</p>



<p class="wp-block-paragraph">Notably, both of the factors that make hardware startups a risky move for first-time entrepreneurs — capital requirements and slow learning curve — are less of a problem for repeat entrepreneurs or industry veterans. People who have amassed capital and experience by starting successful software companies or playing significant roles in existing hardware companies are often the best-positioned to tackle a hardware problem with a new startup.</p>



<p class="wp-block-paragraph">And one-in-a-billion technical experts whose niche knowledge presents a necessary edge in a “hard tech” category may also find that it makes sense for them to start a hardware company. Such cases are especially common in biotech, but can also be seen in fields like nuclear energy and aerospace. For someone who’s spent twenty years becoming the world’s leading expert in a space like gene therapy or nuclear fusion, starting a company in that space can make sense — though that company will still be subject to the risks of high capital requirements and slow cycle times.</p>



<p class="wp-block-paragraph">But for first-time generalist entrepreneurs without experience and capital, it’s hard to beat the advantages of software startups. There’s a reason why Silicon Valley startup culture came into being with the rise of the software industry, and why the vast majority of founders who go from nobodies to big successes make it big initially with software companies. (Dalton and Michael from YC made a <a href="https://www.youtube.com/watch?v=09mXPGVkfVA&amp;list=PLQ-uHSnFig5Nd98Sc9I-kkc0ZWe8peRMC&amp;index=24">great video</a> about this.)</p>



<p class="wp-block-paragraph">The software industry, and the startup ecosystem that’s part of it, are going to undergo big changes in the coming decades under the influence of AI, and may soon present fewer and fewer opportunities for entrepreneurs to make a mark. But until that happens, my advice to aspiring entrepreneurs who are looking for the most promising problems to work on is to strongly prefer founding a software company over a hardware company.</p>]]></content:encoded>
    </item>
    <item>
      <title>The End of the Software Industry</title>
      <link>https://alexkolchinski.com/2023/12/11/the-end-of-the-software-industry/</link>
      <guid isPermaLink="true">https://alexkolchinski.com/2023/12/11/the-end-of-the-software-industry/</guid>
      <pubDate>Mon, 11 Dec 2023 02:01:05 GMT</pubDate>
      <description>This essay is part of a 3-part series: The AI revolution now taking place is poised to change the software industry in fundamental ways, and likely, to shrink the number of people employed in it. Historically, our industry has consisted of smart people painstakingly telling computers how to do specific tasks in the arcane languages […]</description>
      <content:encoded><![CDATA[<p class="wp-block-paragraph" style="line-height:0.5"><strong>This essay is part of a <a href="/2023/12/11/2023-reflections/">3-part series</a>:</strong></p>



<ul class="wp-block-list">
<li><a href="/2023/12/11/the-end-of-knowledge-work/">The End of Knowledge Work</a></li>



<li><a href="/2023/12/11/the-end-of-the-software-industry/">The End of the Software Industry</a></li>



<li><a href="/2023/12/11/founders-beware-hardware/">Founders, Beware Hardware</a></li>
</ul>



<hr class="wp-block-separator has-alpha-channel-opacity is-style-wide" />



<p class="wp-block-paragraph">The AI revolution now taking place is poised to change the software industry in fundamental ways, and likely, to shrink the number of people employed in it. Historically, our industry has consisted of smart people painstakingly telling computers how to do specific tasks in the arcane languages that computers understand. We call this process writing software. AI is upending that model.</p>



<p class="wp-block-paragraph">Already, generative AI tools have substantially increased the productivity of programmers, allowing for more niche or lower value software to be written than previously would have been justifiable economically. But even bigger changes are afoot.</p>



<p class="wp-block-paragraph">We’re moving to a future where AI can conduct arbitrary information-processing tasks based on natural-language instructions from any reasonably intelligent human who understands the problem they’re trying to solve. Manually instructing the computers exactly what to do will no longer be necessary. Essentially, the aims of the no-code movement are coming into focus, though in a different way than might have been anticipated a few years ago —&nbsp;general-purpose AI models are making hand-coded Lego-style no-code tooling obsolete for many use cases.</p>



<p class="wp-block-paragraph">What this means is that there are likely to be a lot fewer programmers soon than there used to be. There may, however, be an increasing number of people doing work along the lines of what product managers and salespeople currently do — understanding and solving customer problems, but with the details of the solutions being realized by the computers themselves instead of by teams of programmers.</p>



<p class="wp-block-paragraph">This parallels the trajectory experienced by many other revolutionary industries, like railroads — after drawing in millions of people in an initial boom, it’s not uncommon for an industry to settle into a relatively stable long-term economic position while relying on fewer and fewer workers to sustain its position.</p>



<figure class="wp-block-image"><img src="https://alexkolchinski.com/wp-content/uploads/2023/12/image-3.png" alt="" class="wp-image-395" /></figure>



<p class="wp-block-paragraph">If the software industry experiences the same dynamic due to more and more of its work being done by AI, this development will also have significant implications for the startup ecosystem. The ability of smart, often young, people with little or no money to start businesses that grow to huge revenues in a decade or less has only really ever been possible in the software industry, which in turn has only existed in its current form for about 50 years. That era is likely to be coming to an end in the near future.</p>



<p class="wp-block-paragraph">This is because a software startup is an organization in which a small group of smart people create new value in the market by building and selling new software — that is, by manually telling computers how to solve a specific human problem. As building software manually becomes less and less necessary, software startups are going to either change in fundamental ways or go extinct entirely. This extinction event, if it happens, will also take the venture capital industry down with the startup ecosystem that feeds it.</p>



<p class="wp-block-paragraph">Of course, how long it takes until AI technology advances to the point where manual software engineering is as antiquated as programming in machine code is today remains to be seen — it could be next year; it could be 50 years from now. But it’s almost certainly coming this century, and even until the full extinction event is complete, the nature of what software is and what it does is going to be constantly changing.</p>



<p class="wp-block-paragraph">This leaves software entrepreneurs, myself included, in a precarious position now. There is more value to be created with software right now than at any point in recent memory, perhaps ever. Trillions of dollars worth of business processes can now be automated, creating huge efficiencies, and trillions more will likely be possible in the next few years. Building and selling software to meet those needs still requires teams of highly-skilled people manually writing code, so there’s a bonanza taking shape in the startup ecosystem — many entrepreneurs are racking up eye-watering revenues with AI products this year, and I expect that to grow substantially in the near future.</p>



<p class="wp-block-paragraph">But many of those startups are likely to be made obsolete almost as quickly as they came into existence. If AI technology advances to the point that a business need can be solved with little to no work on top of general-purpose AI models, startups’ products painstakingly built with manual effort on top of earlier generations of AI technology will quickly lose their value, and those companies will quickly bleed revenue and wither if they haven’t accumulated other advantages in the meantime.</p>



<p class="wp-block-paragraph">Of course, it’s entirely possible that some companies founded in the current AI wave will build more durable assets than quickly-obsolete technology, like owning valuable data or setting themselves up as hard-to-circumvent intermediaries between other companies. But many others are likely to earn huge revenues for a few years by building and selling extremely valuable software, then see those revenues quickly dwindle to zero as their products become trivial to replace due to further advancements in AI.</p>



<p class="wp-block-paragraph">I believe it’s wise for today’s AI entrepreneurs to take this possibility into account. In particular, the possibility of huge revenues up front, followed by a rapid decline, introduces a new dynamic into the cost/benefit calculus of fundraising. In the past, fundraising was usually necessary to get a software product off the ground, because of the significant amount of up-front engineering work normally required to build something that could be sold for meaningful revenue. However, once that revenue was achieved, it was generally durable, because differentiated and valuable business software does not usually become obsolete overnight. Thus, VC funding has been an appropriate source of capital for software companies, which have required risky up-front investment but which have then generated large and durable revenue streams, and commensurate liquidity events, if successful — providing payoffs down the road for both investors and founders.</p>



<p class="wp-block-paragraph">This dynamic may be flipped for many of today’s AI companies. Many previously infeasible or impossible to solve high-value needs in business software are now suddenly solvable in weeks to months with a good team, thanks to recent advancements in AI, and those solutions can quickly generate hundreds of thousands or even millions of dollars in revenue. But that revenue is likely to only last for a few years until the underlying product becomes trivial to replace, unless the company selling the product creates a long-term competitive advantage in ways that are more resilient to rapid advancements in AI.</p>



<p class="wp-block-paragraph">Unfortunately, that dynamic presents significant risks for founders taking venture capital. A founding team able to generate tens or hundreds of millions of dollars in revenue over 5-10 years with a product that then vanishes in a puff of smoke will see very different outcomes depending on whether or not it takes venture capital. Without external investment, those revenues, net of costs, will be the founders’ to keep. With external investment, the founders will be able to take a relatively meager salary, but will then be expected to plow those revenues into future growth, which may not be available if the product direction is obsolete. In this way, raising even small amounts of venture capital can turn a multimillion-dollar outcome into a zero for the founders involved.</p>



<p class="wp-block-paragraph">My conclusion is that for teams building in today’s AI landscape, it may be wise to forego fundraising entirely, at least at first, and instead to focus on generating as much revenue as possible, as early as possible. In many cases, it may be possible to get to millions of dollars in revenues with a small, self-funded team, then swing for the fences and attempt to create a huge and durable company from there. But if that proves not to be possible, at least the founders of such a company will have a sizable, if temporary, profit stream to hedge the risk of a dead end.</p>



<p class="wp-block-paragraph">And for the software founders involved, being able to guarantee personal financial security may be more important now than ever, as our ability to generate economic value is likely to decline or even vanish in the decades to come as AI becomes better at building software than we are. Now is the time to gather our rosebuds while we may — and to reduce the risk of a zero outcome if possible.</p>]]></content:encoded>
    </item>
    <item>
      <title>The End of Knowledge Work</title>
      <link>https://alexkolchinski.com/2023/12/11/the-end-of-knowledge-work/</link>
      <guid isPermaLink="true">https://alexkolchinski.com/2023/12/11/the-end-of-knowledge-work/</guid>
      <pubDate>Mon, 11 Dec 2023 01:59:36 GMT</pubDate>
      <description>This essay is part of a 3-part series: For almost all of human history, the vast majority of people had to rely on their brains and brawn to survive, through hunting and gathering and then agriculture. Animal power provided some assistance, but things started to change much more fundamentally once people learned to harness non-biological […]</description>
      <content:encoded><![CDATA[<p class="wp-block-paragraph" style="line-height:0.5"><strong>This essay is part of a <a href="/2023/12/11/2023-reflections/">3-part series</a>:</strong></p>



<ul class="wp-block-list">
<li><a href="/2023/12/11/the-end-of-knowledge-work/">The End of Knowledge Work</a></li>



<li><a href="/2023/12/11/the-end-of-the-software-industry/">The End of the Software Industry</a></li>



<li><a href="/2023/12/11/founders-beware-hardware/">Founders, Beware Hardware</a></li>
</ul>



<hr class="wp-block-separator has-alpha-channel-opacity is-style-wide" />



<p class="wp-block-paragraph">For almost all of human history, the vast majority of people had to rely on their brains and brawn to survive, through hunting and gathering and then agriculture. Animal power provided some assistance, but things started to change much more fundamentally once people learned to harness non-biological energy: wind, water, and then most impactfully steam.</p>



<p class="wp-block-paragraph">With the advent of more plentiful energy and the development of more capable mechanical technology, the world changed. Jobs that used to be done with muscle power became easier to do with machines, and the roles of people in society changed as well.</p>



<p class="wp-block-paragraph">Up until the 19th century, most workers’ jobs were simply to make food. But over the following century, agriculture became highly mechanized, and the fraction of people in industrialized countries that worked as full-time farmers plummeted:</p>



<figure class="wp-block-image"><img src="https://alexkolchinski.com/wp-content/uploads/2023/12/image-2.png" alt="" class="wp-image-392" /></figure>



<p class="wp-block-paragraph">At first, this agricultural revolution freed people to become other kinds of manual workers — especially in manufacturing and transportation, where they made the wide variety of physical implements needed for the new industrialized economy, and moved them to where they were needed.</p>



<p class="wp-block-paragraph">But over time, increasingly capable technology has caused the non-agricultural blue-collar part of the economy to require less and less labor as well. A factory that took 1000 people to run might now run with 30 technicians keeping an eye on the machines doing most of the work; a train that used to take 10 people to operate might now run with one.&nbsp;</p>



<figure class="wp-block-image"><img src="https://alexkolchinski.com/wp-content/uploads/2023/12/image-1.png" alt="baker buffie blue collar 2016 02 21 1" class="wp-image-391" /></figure>



<p class="wp-block-paragraph">Instead, in the industrialized world, people have increasingly become economically useful for their brains rather than for their brawn. A majority of US workers are now in white-collar fields, where they are paid for their knowledge and intellectual skills rather than for their strength and dexterity.</p>



<figure class="wp-block-image size-large is-resized"><img loading="lazy" width="1024" height="646" src="https://alexkolchinski.com/wp-content/uploads/2023/12/image.png" alt="" class="wp-image-389" style="width:538px;height:auto" /></figure>



<p class="wp-block-paragraph">This process has had a wide range of effects on society. Physical strength and skill used to be a necessity for survival for most people. Now, an unusually strong American is as likely to be a hobbyist weightlifter as a manual laborer. Muscles used to be a tool; increasingly, they’ve become an ornament.</p>



<p class="wp-block-paragraph">Instead, in developed countries like America, most people’s brains have become their most economically valuable assets. We’re becoming an economy where our interactions with the physical world are carried out by increasingly capable tools that require less and less physical effort from us — but where the management of an increasingly complicated world rests more and more on highly capable human brains.</p>



<p class="wp-block-paragraph">Even our identities, to a significant degree, are tied up with our intelligence. We are homo sapiens, the wise human — we ascribe our preeminence as a species to our ability to understand and analyze the world around us.</p>



<p class="wp-block-paragraph">It is that aspect of our existence that is now starting to change in the same way that our physical interactions with the world have been changing for the last 200 years.</p>



<p class="wp-block-paragraph">Much like with draft animals before the Industrial Revolution, we’ve had some external assistance for our minds before now. Technology for communicating knowledge, from books to the Internet, has become more and more advanced over the centuries, giving us a longer and longer lever for our intellectual efforts. And computational technology, from abaci to today’s computers, has become more and more capable of doing routine but voluminous informational work quickly for us.</p>



<p class="wp-block-paragraph">This has already changed the composition of the white-collar labor force: human “calculators” have been replaced by spreadsheets, and draftsmen by CAD programs. But bigger changes are now afoot.</p>



<p class="wp-block-paragraph">We’ve reached a point in the development of artificial intelligence where its capabilities are extending into previously unimaginable territory. Computers used to need precise instructions for every task they could carry out, and they could only accept those instructions in their own arcane languages. Over the past year, this has changed — computers can now accept instructions in human language, written or spoken, and carry out those instructions successfully for even quite complex tasks.</p>



<p class="wp-block-paragraph">Those capabilities are still far from perfect, but even if the abilities of AI were to be frozen at today’s level, white-collar work being performed by the full-time equivalent of many millions of people can, and will, be automated in the years to come using the technology that’s already available.</p>



<p class="wp-block-paragraph">And future advances are inevitable as well. To date, most improvements in AI capabilities have come from deploying increasingly large amounts of computational power to train AI models on increasingly large amounts of data. Roughly speaking, every meaningful increase in AI capabilities has resulted from multiplying the amount of data and computational power used to train AI models by a factor of 10.&nbsp;</p>



<p class="wp-block-paragraph">It’s likely that this approach will continue to bear fruit for at least one more “10x” cycle, which will likely take one to several more years. After that, data availability is likely to become a bottleneck — at that point, the latest AI models will have been trained on most extant codified human knowledge (books, the Internet, etc.) Feeding another “10x” cycle by synthesizing fake but realistic data, or gathering much more data from the real world, appears possible, but will likely be much slower than the current approach of using data that’s already available.</p>



<p class="wp-block-paragraph">But in theory, it should also be possible to massively improve the data efficiency of training AI systems. Human brains have about 100 billion neurons, and humans can acquire a respectable education by learning information equivalent to the contents of a few hundred to a few thousand books. But state-of-the-art AI models, still far less capable than human brains, require many GPUs to train, and each GPU contains roughly as many transistors as the human brain contains neurons. And those state-of-the art models are being fed the equivalent of millions or even billions of books, far more than any human could ever read.</p>



<p class="wp-block-paragraph">Thus, it seems likely that advances are possible both in the architectures and training methods we use to make AI models, as well as in the silicon hardware we train them on.</p>



<p class="wp-block-paragraph">So, if and when significant advances in training efficiency happen, we’re going to be able to train much more capable AI models using the data, and likely the hardware, that’s already available. And soon enough — whether it’s in one year or twenty — the capabilities of AI in all analytical tasks are going to exceed those of most or all humans.</p>



<p class="wp-block-paragraph">The consequences of this revolution — even extrapolating the capabilities of the technology that’s already available today, let alone that of more advanced systems — are going to be enormous, and are going to happen much more quickly than the Industrial Revolution. Unlike physical machines, AI technology can be deployed instantly to the whole world, meaning each advancement can be adopted as quickly as people can figure out how to use it, rather than being limited by the speed at which physical machines can be built and deployed.</p>



<p class="wp-block-paragraph">What this implies is that much of the white-collar work that currently consists of analytical tasks — essentially, manipulating information&nbsp;— is going to be automated in the blink of an eye. A million people might work in data entry today — next year, that number might be zero. In a few years, when the capabilities of the technology advance, chemists or programmers might be similarly affected, and the process will continue until humans are only doing a small fraction of the analytical work that we do today.</p>



<p class="wp-block-paragraph">In reality, in most lines of work, the AI won’t fully automate an entire profession —&nbsp;rather, it’ll reduce the amount of human effort required in that line of work by 10x, or 100x, or 1000x, just like what physical automation achieved in factories.</p>



<p class="wp-block-paragraph">As a consequence, we’re going to see the raw information-processing abilities of human brains become less and less economically valuable. Technical specialists like programmers and traders, who work with self-contained purely-informational tasks, are going to see some of the biggest changes as soon as AI’s abilities exceed theirs. But most white-collar jobs contain a significant component of information processing, and are going to see that quickly handed over to the machines.</p>



<p class="wp-block-paragraph">Interestingly, it seems that interpersonal relationship-oriented work is likely to be much less affected. Jobs ranging from bartender to bond salesperson rely heavily on interpersonal relationship-building, and much of the value people create in those jobs is dependent on them being people and not machines. In addition, high-end jobs that have as much to do with brokering ownership of valuable assets, power, and influence&nbsp;— think politicians, investment bankers, and influencers — are going to continue to be done by the people who care about accumulating that ownership, power, and influence, although those people are going to be increasingly assisted in their work by powerful AI, just as they are currently assisted by human workers.</p>



<p class="wp-block-paragraph">Physical work will also be initially unaffected by the current revolution in AI — rescuing someone from a burning building or making a bed is still only possible with human hands. But the current AI revolution is increasingly looking like it’s going to unlock advances in robotics that will have a significant impact on the automation of physical work as well.</p>



<p class="wp-block-paragraph">Currently, most physical-work automation is done by highly-specialized machines — walk into a reasonably-modern factory and you’ll see a number of large and expensive machines custom-tailored for jobs like bending rods or forming cans. This kind of special-purpose automation is only feasible for tasks where huge scale can be achieved in a single place — easier for tasks like manufacturing and agriculture than for tasks like house-cleaning and food-service, the need for which is naturally dispersed throughout the physical world and cannot be performed in a single place.&nbsp;</p>



<p class="wp-block-paragraph">That said, mechanical technology is already advanced enough to make it possible to automate many of these highly-dispersed physical tasks that are still being done mostly by humans. The missing factor standing in the way of automating those tasks has rather been the intelligence available to the machines — robots have just not been smart enough to be able to do things like reliably navigate a house and manipulate a wide variety of objects.</p>



<p class="wp-block-paragraph">That is all likely to change soon as the capabilities of AI continue to advance, and while the consequences will take longer to be realized in physical-world automation than with knowledge work, as they involve the manufacture and deployment of physical robots, the process will likely be faster than people expect.&nbsp;</p>



<p class="wp-block-paragraph">This is because AI is going to unlock the wide use of a new type of automation: automation done by general-purpose hardware, whether that hardware takes the shape of humanoid robots, robot arms mounted on quadrupedal platforms, or something else entirely. And when one-size-fits-all robotic platforms powered by AI become broadly capable, they will be much cheaper and faster to manufacture and deploy than a heterogenous variety of specialized machines. This is exactly the same process thanks to which cars are now so cheap and ubiquitous — when you’re building a billion of something, you get a lot of efficiencies of scale. And when you’re building a billion general-purpose robots, which can do a wide variety of work thanks to AI, you can automate a lot of that work very quickly.</p>



<p class="wp-block-paragraph">In this way, much of the work being done by people today — physical as well as analytical — is going to be increasingly done by machines. This process has already started, and is likely to accelerate over the next decade or two.&nbsp;</p>



<p class="wp-block-paragraph">Where that leaves us humans in the economy is another question. It’s possible that a parallel process to the last 200 years may take place. Since industrialization, the physical abilities of people have become steadily less in-demand as machines have taken our places in manipulating the physical world, but through that process, the demand for our brains has only increased, causing more and more people to make a living by knowledge work as fewer and fewer make a living by physical work.&nbsp;</p>



<p class="wp-block-paragraph">In a similar way, we may see less and less economic demand for our brains as analytical machines over the next decade or two, but to see more and more demand for our ability to relate to each other. In that kind of world, most of us will have jobs talking to each other about our needs and looking for ways to solve them — everyone a salesperson or therapist. In a world like that, brainpower will still be important, but more as a way to relate to each other than as a tool to solve complex puzzles. Incidentally, one theory of human evolution states that we evolved big brains as a result of social selection pressures that favored the survival and reproduction of those of us who were better at relating to other people — after a few hundred years of using our brains in more-analytical ways than they evolved for, we may return to a world in which our brains become useful primarily for the interpersonal tasks that they were honed by evolution to do.</p>



<p class="wp-block-paragraph">But it’s also possible that AI will do such a good job at facilitating and even replacing human interactions and relationships in economic contexts (sales, customer support, etc.) that the number of humans needed for that kind of work will be far less than the number of people who will need to earn a living. In one sense, that’s a scary world — most of us will no longer be able to sustain a decent standard of living through the work we can do. But if that’s the way the world develops, the only politically feasible outcome I foresee, whether in democracies or dictatorships, is a welfare state where the work done to fulfill human needs is mostly done autonomously, and people are able to enjoy a high standard of living without working. Instead, most people would be able to spend their time on things that might not have paid the bills before, whether taking care of their families, competing at sports, traveling, painting, or any number of other things that mostly exist outside of the market economy today.</p>



<p class="wp-block-paragraph">Which of those two scenarios comes to pass, or whether another one entirely will transpire, only time will tell. And whether the whittling-down of the need for knowledge work by humans takes half a decade or half a century will remain to be seen as well. But our need to work with our minds is about to go through the same winnowing process as our need to work with our bodies has experienced for the past 200 years — and it remains to be seen what the consequences will be.</p>]]></content:encoded>
    </item>
    <item>
      <title>How to talk to ChatGPT through Siri</title>
      <link>https://alexkolchinski.com/2023/03/01/how-to-talk-to-chatgpt-through-siri/</link>
      <guid isPermaLink="true">https://alexkolchinski.com/2023/03/01/how-to-talk-to-chatgpt-through-siri/</guid>
      <pubDate>Wed, 01 Mar 2023 21:07:22 GMT</pubDate>
      <description>Recently, I wrote the post How to: Talk to GPT-3 Through Siri, describing how to significantly upgrade Siri using OpenAI’s recent davinci-003 model. The iOS shortcut in that post is already a big upgrade to Siri, but davinci-003 isn’t as capable as OpenAI’s latest models, which are also what is powering ChatGPT behind the scenes. […]</description>
      <content:encoded><![CDATA[<p class="wp-block-paragraph">Recently, I wrote the post <a href="/2023/02/03/how-to-talk-to-gpt-3-through-siri/">How to: Talk to GPT-3 Through Siri</a>, describing how to significantly upgrade Siri using OpenAI&#8217;s recent davinci-003 model.</p>



<p class="wp-block-paragraph">The iOS shortcut in that post is already a big upgrade to Siri, but davinci-003 isn&#8217;t as capable as OpenAI&#8217;s latest models, which are also what is powering ChatGPT behind the scenes.</p>



<p class="wp-block-paragraph">Until now, those models weren&#8217;t available for API access, but today, OpenAI opened API access to their gpt-3.5-turbo model, and I&#8217;ve updated the shortcut so you can now talk to the equivalent of ChatGPT directly through Siri.</p>



<p class="wp-block-paragraph">gpt-3.5-turbo is about 10x cheaper to use than davinci-003, and seems to return better answers for some questions, but worse answers for others. Your call on which to use! I&#8217;ve switched to gpt-3.5-turbo, personally.</p>



<p class="wp-block-paragraph"><strong><a href="https://www.icloud.com/shortcuts/699314183d0a49489bf25087957c7a89">You can download the shortcut here</a></strong> – to start using it, you&#8217;ll just need to input your OpenAI API key. Then, talk to ChatGPT by saying &#8220;Hey Siri, GPT Mode&#8221; (or whatever you rename the shortcut to), then your question.</p>



<p class="wp-block-paragraph">If you want to have Siri consistently read the responses out loud, it&#8217;s also best to change Siri’s settings in Settings-&gt;Accessibility-&gt;Siri-&gt;Spoken Responses to “Prefer Spoken Responses”:</p>



<figure class="wp-block-image is-resized"><img loading="lazy" src="https://alexkolchinski.com/wp-content/uploads/2023/02/image.png" alt="" class="wp-image-325" width="236" height="512" /></figure>



<p class="wp-block-paragraph">If you want more detailed instructions on how to get this working, please see <a href="/2023/02/03/how-to-talk-to-gpt-3-through-siri/">my previous post</a>.</p>



<p class="wp-block-paragraph">PS – No pressure, but if you&#8217;ve found this shortcut useful, I&#8217;d appreciate it if you <a href="https://www.buymeacoffee.com/kolch">buy me a coffee</a>!</p>]]></content:encoded>
    </item>
    <item>
      <title>How To: Talk to GPT-3 through Siri</title>
      <link>https://alexkolchinski.com/2023/02/03/how-to-talk-to-gpt-3-through-siri/</link>
      <guid isPermaLink="true">https://alexkolchinski.com/2023/02/03/how-to-talk-to-gpt-3-through-siri/</guid>
      <pubDate>Fri, 03 Feb 2023 18:54:44 GMT</pubDate>
      <description>Note: it’s now possible to talk to a newer OpenAI model (gpt-3.5-turbo) through Siri – if you want to use the newer, and much cheaper, version, see my updated post here. The new version seems to do better on some questions but worse on others. Like many others, I’ve been incredibly impressed with OpenAI’s ChatGPT […]</description>
      <content:encoded><![CDATA[<p class="wp-block-paragraph"><strong>Note: it&#8217;s now possible to talk to a newer OpenAI model (gpt-3.5-turbo) through Siri – if you want to use the newer, and much cheaper, version, see my updated post <a href="/2023/03/01/how-to-talk-to-chatgpt-through-siri/">here</a>.</strong> <strong>The new version seems to do better on some questions but worse on others.</strong></p>



<p class="wp-block-paragraph">Like many others, I&#8217;ve been incredibly impressed with OpenAI&#8217;s <a href="https://openai.com/blog/chatgpt/">ChatGPT</a> and how far language models have come since I was working on natural language processing research a few years ago.</p>



<p class="wp-block-paragraph">But, also like many others, I&#8217;ve been regularly frustrated with Apple&#8217;s Siri and how it often fails to give useful answers to even the most basic of questions. This week, after a few too many unsatisfying Siri answers in a row, I started to wonder if it would be possible to solve the problem once and for all by querying ChatGPT directly through Siri.</p>



<p class="wp-block-paragraph">It turned out that while the ChatGPT API isn&#8217;t officially available yet to pose questions to programmatically, the recent GPT-3 model is, and it&#8217;s also very powerful.</p>



<p class="wp-block-paragraph">It also turned out that several people have written Siri shortcuts to interface with GPT-3, but I had a couple of issues with them when I tried them out:</p>



<ul class="wp-block-list">
<li>Siri wouldn&#8217;t always read the answers out loud</li>



<li>The answers often started with a stray question mark, which was distracting when Siri did read the answer out loud, starting with &#8220;Question mark&#8230;&#8221;</li>
</ul>



<p class="wp-block-paragraph">I tried fixing the first issue by making a shortcut that included steps to explicitly read the answer out loud, but it turned out that a far simpler solution was just to change Siri&#8217;s settings in Settings-&gt;Accessibility-&gt;Siri-&gt;Spoken Responses to &#8220;Prefer Spoken Responses&#8221;:</p>



<figure class="wp-block-image size-large"><img loading="lazy" width="472" height="1023" src="https://alexkolchinski.com/wp-content/uploads/2023/02/image.png" alt="" class="wp-image-325" /></figure>



<p class="wp-block-paragraph">Make sure you turn on &#8220;Prefer Spoken Responses&#8221; if you want to use this shortcut and have Siri read the answers out loud to you!</p>



<p class="wp-block-paragraph">The stray leading question marks were an easier fix – I just modified an existing shortcut with a step that strips them out.</p>



<p class="wp-block-paragraph">With those two fixes, the shortcut started working very seamlessly – I can now tell my phone &#8220;Hey Siri, GPT Mode&#8221;, then a question, and quickly get a response from GPT-3 read back to me by Siri.</p>



<p class="wp-block-paragraph"><a href="https://www.icloud.com/shortcuts/b10d3d361a3f48428a2ed8fe729dc4fa"><strong>You can download the Siri shortcut here</strong></a> to add it to your phone (and you can see the original shortcut that I modified it from <a href="https://www.mobilespoon.net/2023/01/how-to-activate-chatgpt-with-siri-and-save-response.html">here</a>).</p>



<p class="wp-block-paragraph">The shortcut itself is free to use; you&#8217;ll just need to create an OpenAI account <a href="https://openai.com/api/">here</a>, create an OpenAI API key and paste it into the text field in the shortcut that says &#8220;Replace this with your OpenAI API key!&#8221; You can see more detailed instructions for how to do this in <a href="https://www.macobserver.com/tips/how-to/integrate-chatgpt-siri/">this article explaining how to use a similar shortcut.</a> Right now, OpenAI is offering three months of API credits for free when you sign up for an account, and they don&#8217;t charge much once you run out of your free credits.</p>



<p class="wp-block-paragraph">Once you&#8217;ve installed the shortcut onto your phone and pasted your OpenAI API key into it, it&#8217;s ready to use! You can use it under the name I gave it (&#8220;GPT Mode&#8221;) or rename it to your taste, to anything that doesn&#8217;t conflict with Siri&#8217;s existing commands.</p>



<p class="wp-block-paragraph">To use the shortcut, say &#8220;Hey Siri&#8221;, then &#8220;GPT Mode&#8221; (or whatever you renamed the shortcut to), then say whatever you want to ask GPT-3. Siri will then, after a delay, display the answer and read it out loud. </p>



<p class="wp-block-paragraph">Much like with ChatGPT, the answers aren&#8217;t always right, but they usually are, and I&#8217;ve found that the ability to ask my phone complex questions and get reasonable answers back is incredibly useful as long as I keep the imperfections of GPT-3 in mind.</p>



<p class="wp-block-paragraph">Hopefully, the day is coming soon when Siri will have this kind of functionality built in natively! In the meantime, there&#8217;s this Siri shortcut.</p>



<p class="wp-block-paragraph">PS – No pressure, but if you&#8217;ve found this shortcut useful, I&#8217;d appreciate it if you <a href="https://www.buymeacoffee.com/kolch">buy me a coffee</a>!</p>]]></content:encoded>
    </item>
    <item>
      <title>Crypto is an Unproductive Bubble</title>
      <link>https://alexkolchinski.com/2022/03/18/crypto-is-an-unproductive-bubble/</link>
      <guid isPermaLink="true">https://alexkolchinski.com/2022/03/18/crypto-is-an-unproductive-bubble/</guid>
      <pubDate>Fri, 18 Mar 2022 23:06:03 GMT</pubDate>
      <description>This post sparked a great conversation on Hacker News! See the comments here.Another update: Scott Alexander just (12/8/22) wrote the best counterpoint to this that I’ve read – see it here “The four most expensive words in investing are: ‘This time it’s different.’” John Templeton I’m writing this essay not so much to convince anyone […]</description>
      <content:encoded><![CDATA[<p class="wp-block-paragraph"><em>This post sparked a great conversation on Hacker News! See the comments <a href="https://news.ycombinator.com/item?id=30728856">here</a>.</em><br><em>Another update: Scott Alexander just (12/8/22) wrote the best counterpoint to this that I&#8217;ve read – see it <a href="https://astralcodexten.substack.com/p/why-im-less-than-infinitely-hostile?utm_source=substack&amp;utm_medium=email">here</a></em></p>



<blockquote class="wp-block-quote is-layout-flow wp-block-quote-is-layout-flow">
<p class="wp-block-paragraph">&#8220;The four most expensive words in investing are: &#8216;This time it&#8217;s different.'&#8221;</p>
<cite>John Templeton</cite></blockquote>



<p class="wp-block-paragraph">I&#8217;m writing this essay not so much to convince anyone of my point of view as to outline my current thinking around cryptocurrencies and the mania surrounding them. Because this is a controversial subject over which smart people disagree, I&#8217;m making these claims in public so that I might be proven either right or wrong later.</p>



<p class="wp-block-paragraph">If you disagree with any of my claims here, or have more to add, I&#8217;d appreciate you leaving a comment with your thoughts, both for my benefit and for that of other readers.</p>



<h3 class="wp-block-heading">Claim 1: Crypto is a Bubble (Confidence: High)</h3>



<p class="wp-block-paragraph">The hallmark of a bubble is people buying an asset primarily in the hopes of future appreciation driven by other buyers, rather than because of any notion of its long-term value. Cryptocurrencies have satisfied this criterion starting a number of years ago, and from what I&#8217;ve seen continue to do so. The trouble with bubbles is that once the supply of hopeful buyers runs out, prices stop rising and start falling as fear replaces greed. I predict that this will happen to all cryptocurrencies, including Bitcoin and Ethereum, within a decade. This goes doubly for double-bubble assets like NFTs.</p>



<p class="wp-block-paragraph">Crypto boosters claim that cryptocurrency has long-term value as a digital currency, but I disagree for the reasons outlined in Claim 3.</p>



<h3 class="wp-block-heading">Claim 2: Blockchain technology has no non-monetary applications (Confidence: High)</h3>



<p class="wp-block-paragraph">While this mania seems to have calmed down as of late, I&#8217;ve seen a number of claims that blockchain technology – i.e. the ability for a network of people to maintain a distributed ledger without trusting a central authority or authorities – unlocks non-monetary uses, e.g. supply chain transparency. However, any use of blockchain technology in which the ledger is not self-contained, but is instead tied to the physical world, has the fatal flaw that there can be no trust-free link between the ledger and the physical world. If you don&#8217;t trust anybody to enter a delivery of wheat into your ledger, then your supply-chain ledger will be empty, but if you do trust certain people, then you are relying on the honesty of specific people or on other safeguards outside the ledger itself, undermining the purpose of a distributed ledger. </p>



<h3 class="wp-block-heading">Claim 3: Future monetary use of blockchain technology will be minor (Confidence: Medium)</h3>



<p class="wp-block-paragraph">This is where it gets more tricky. Blockchain technology does unlock an interesting use case, which is the ability of participants in a network to maintain a ledger, or record of holdings, without trusting a central authority. In theory, this approach could replace fiat currency, which is controlled and tracked by trusted authorities like banks. However, I see two major roadblocks to the widespread adoption of cryptocurrencies:</p>



<ol class="wp-block-list">
<li><strong>Inefficiency</strong> – Requiring every participant in a network to interact with a distributed ledger imposes substantial transaction costs which make cryptocurrency transactions more expensive than transactions in fiat currency, where the network is supported by trusted central authorities like banks and the government. To my knowledge, nobody has yet found a way to correct this flaw and it&#8217;s likely a feature of the way distributed ledgers work. If that&#8217;s the case, then cryptocurrency adoption will be limited to marginal use cases where fiat currency is impractical, e.g. online black-market transactions, since people prefer paying lower rather than higher transaction fees.</li>



<li><strong>Control of monetary policy</strong> – When the world was on the gold standard and central banks had limited ability to control the money supply, extreme and frequent booms and busts were a major feature of the economy and a constant detriment to people&#8217;s lives and livelihoods. Fiat currencies administered by competent central bankers have the attractive feature that their governance can be used as a tool to cool excesses and calm panics when the credit cycle goes through its inevitable revolutions. Cryptocurrencies, to my understanding, dispense with this notion of central control on purpose, and their adoption as the primary currency in an economy would subject that economy to pre-fiat-era booms and busts or worse. </li>
</ol>



<p class="wp-block-paragraph">Unless both of these factors are somehow addressed, I predict that cryptocurrencies will play a marginal role in the future, if they exist at all. Even if they continue to grow in value as more speculators pile in, their usage as currency will be limited unless the inefficiency problem is solved. If that does somehow happen as well, we may see widespread use of cryptocurrencies in at least some countries, until a reckoning with the subsequent booms-and-bust dynamics, which may take decades, forces a return to fiat currency or its equivalent. (This could take the form of something which is a cryptocurrency in name only, with a &#8220;ledger&#8221; maintained by the government and banks.)</p>



<p class="wp-block-paragraph">I also believe that it&#8217;s unlikely we&#8217;ll get even that far, as governments will not be keen to lose control over monetary policy. Authoritarian governments are likely to restrict or ban cryptocurrencies if they get much bigger than they are now, and even democratic governments will have to weigh voters&#8217; enthusiasm for cryptocurrencies with the importance of being able to control monetary policy.</p>



<p class="wp-block-paragraph">Some bubbles are caused by over-enthusiasm over genuinely innovative and productive assets, like the Railway Mania of the 1840s or the Internet Bubble of the 1990s. When those bubbles pop, they leave behind large amounts of investment in assets like rail and software that can be used for productive purposes later. (For more on this, see Carlota Perez&#8217;s <em>Technological Innovations and Financial Capital</em>, or a summary.) Other bubbles are triggered by enthusiasm over assets which have much less enduring value, like Tulip Mania, the South Sea Bubble, and Beanie Babies. It&#8217;s my prediction that the cryptocurrency bubble will turn out to have been another one of those unproductive bubbles, fueled by low interest rates and speculative enthusiasm from a wide base of retail investors rather than by a genuine advance in financial technology.</p>



<p class="wp-block-paragraph"></p>]]></content:encoded>
    </item>
    <item>
      <title>Regulating Resiliency in Supply Chains</title>
      <link>https://alexkolchinski.com/2021/10/08/regulating-resiliency-in-supply-chains/</link>
      <guid isPermaLink="true">https://alexkolchinski.com/2021/10/08/regulating-resiliency-in-supply-chains/</guid>
      <pubDate>Fri, 08 Oct 2021 21:30:35 GMT</pubDate>
      <description>COVID has exposed just how vulnerable to shocks the systems we depend on for physical goods are. Just about every good – from lumber to computer chips – has been subject to shortages, and each shortage has reverberated downstream through supply chains to cause shortages in many other goods, as with the millions of cars […]</description>
      <content:encoded><![CDATA[<p class="wp-block-paragraph">COVID has exposed just how vulnerable to shocks the systems we depend on for physical goods are. Just about every good – from <a href="https://cnr.ncsu.edu/news/2021/05/lumber-shortage/">lumber</a> to <a href="https://en.wikipedia.org/wiki/2020%E2%80%932021_global_chip_shortage">computer chips</a> – has been subject to shortages, and each shortage has reverberated downstream through supply chains to cause shortages in many other goods, as with the millions of cars that aren’t being made because of chip shortages. All these shortages are having a meaningful effect on consumers’ quality of life and on firms’ ability to produce goods and conduct R&amp;D – both a short-term and long-term hit to our economy and well-being.</p>



<p class="wp-block-paragraph">Why are physical supply chains so fragile? Part of the reason is just that they’re so complex. A firm making a physical good likely sources components and tools from dozens, if not hundreds or thousands, of upstream suppliers. Often, many of those upstream suppliers are in a different country than the firm in question. It only takes an issue with one of those upstream suppliers – or with the ability to transport one of their products – to throw a firm’s operations off-kilter. In some cases, a shortage of a component or tool means scrambling to find an alternative supplier, but in some cases – as with chips – there may be no alternative, and the firm’s operations become limited by the quantity that a single supplier is able to deliver. Either way, production slows or stops, having cascading effects on the downstream customers that depend on the firm in question’s products.</p>



<p class="wp-block-paragraph">So, any link in a supply chain can cause disproportionate harm to the entire supply chain relative to its own size. But in many cases, the links – firms producing physical goods – are actually incentivized to be efficient to the point of fragility. Many physical goods, whether toilet paper or lumber, are essentially fungible commodities in near-perfectly competitive markets, meaning firms producing those goods must eke out every possible efficiency to defend their already-thin margins. This means running production equipment 24/7 at full capacity and maintaining stocks of inputs and inventory that are as small as possible to minimize capital outlay and storage costs.&nbsp;</p>



<p class="wp-block-paragraph">This is, of course, a fragile way to do things, but it’s often the only economically-viable one, and firms are forced to risk being unable to meet a once-a-decade crisis in order to preserve day-to-day profitability. The trouble is that when a firm is forced to slow down or shut down by such a crisis, only a small fraction of the economic costs of the shutdown are borne by the firm itself – many more are borne by the supply chain and economy within which the firm operates because of goods left unproduced and unused. (a negative externality, in economic terms)</p>



<p class="wp-block-paragraph">There are also cases where a physical good, rather than being produced by many near-identical firms, is produced by a single firm, or only a few. This is the case with both computer chips and chip production equipment, with top-end supply dominated by <a href="https://en.wikipedia.org/wiki/TSMC">TSMC in Taiwan</a> and <a href="https://en.wikipedia.org/wiki/ASML_Holding">ASML Holding in the Netherlands</a> respectively. Just as with the paper-thin operational slack in commodity goods, the concentration of production in goods like chips makes sense: for goods that are extremely expensive to produce in terms of fixed costs of intellectual property and physical infrastructure, a natural monopoly or oligopoly is the expected result.&nbsp;</p>



<p class="wp-block-paragraph">The trouble is that when only one or a few firms are at a critical supply chain chokepoint, a single event can cause a worldwide disruption, as with the <a href="https://en.wikipedia.org/wiki/2011_Thailand_floods#Damages_to_industrial_estates_and_global_supply_shortages">2011 floods in Thailand</a> that upended the hard drive industry. Moreover, firms operating expensive capital equipment are also incentivized to run at 100% capacity just like a toilet paper producer, meaning that there’s little room to absorb spikes in demand – as with TSMC struggling to meet chip demand during the COVID recovery.</p>



<p class="wp-block-paragraph">These kinds of vulnerabilities might seem inevitable, but there is room to reduce them with regulation.&nbsp;</p>



<p class="wp-block-paragraph">Another industry, finance, is concerned with moving money rather than physical goods, but has similar vulnerabilities – individual financial firms are also incentivized to operate in risky ways that make sense for each of them individually but not for the system as a whole, since any firm’s failure can be magnified through the whole financial system and real economy.</p>



<p class="wp-block-paragraph">And indeed, the world of finance used to suffer from <a href="https://en.wikipedia.org/wiki/List_of_economic_expansions_in_the_United_States">catastrophic booms and busts every decade or so</a>, impacting the real economy of goods and services in major and painful ways. However, regulation in the last century has reduced, though not eliminated, the frequency of cascading failures in the financial system. A similar approach might be useful for physical supply chains.</p>



<p class="wp-block-paragraph">Of course, there are limitations to the analogy. In a financial crisis, the government can inject money it creates out of thin air to improve liquidity. Unfortunately, the Fed is not able to materialize computer chips or toilet paper out of thin air in the same way that it can manifest new dollars. So, regulation for physical supply chains should be aimed at making them more resilient before a crisis strikes.</p>



<p class="wp-block-paragraph">There are a few ways in which this can be done.</p>



<p class="wp-block-paragraph">One is the “too big to fail” principle, or ensuring that no single firm is so big, or so much the exclusive supplier of some critical good, that its failure would cause the whole system to crumble. In addition to sheer scale, in physical supply chains, there is also the importance of international borders – even if there are many suppliers of a good, but all are located in a single foreign country, then it is that country’s government’s willingness to export that good that becomes a single failure point.</p>



<p class="wp-block-paragraph">Thankfully, the US is already starting to wake up to this concern. Projects like the new Intel plant and new TSMC plant, both to be located in the US, are reducing vulnerability in chip supply chains. Similar projects, supported by the state both politically and with investment capital, could reduce vulnerabilities in other key supply chain chokepoints.</p>



<p class="wp-block-paragraph">Another issue in physical supply chains is somewhat analogous to bank reserve requirements. Banks are incentivized to keep relatively little money on hand – what is not lent out is not earning interest. But low reserves increase the fragility of the whole financial system, since banks with low reserves are more vulnerable than those with ample ones, and a single bank failure can cascade through the whole system.</p>



<p class="wp-block-paragraph">Similarly, firms producing physical goods in highly-competitive spaces are incentivized to maintain very low inventories of both inputs and finished goods, and very low if any excess production capacity. But just like with banks with small reserves, goods-producing firms with small inventories and small excess production capacities are vulnerable to crises and cannot meet bursts in demand. Just like with banks, this makes the whole system more fragile.</p>



<p class="wp-block-paragraph">One can imagine the government stepping in and instituting similar “reserve requirements” for goods-producing firms as for banks. For example, firms over a certain size might be required to hold a larger inventory buffer of inputs and outputs on hand than would be economically rational, and to be subsidized by the government for the additional costs incurred, just like banks are paid interest by the government on their reserves.&nbsp;</p>



<p class="wp-block-paragraph">In addition, the government might incentivize or require firms over a certain size to maintain excess production capacity. This could take the shape of a Federal program where firms were periodically required on short notice to produce an extra 20% of their annual production of a good within some number of months, which would then be bought at above-market rates by the government. That’s one example, but I’m sure those more familiar with industrial regulation could imagine even better ones.</p>



<p class="wp-block-paragraph">Regulating companies producing physical goods is no easy matter, and any regulation that improves supply chain resiliency would impose costs on taxpayers. But if done properly, it could also reduce our vulnerability to future supply chain shocks and serve as a form of worthwhile insurance to our economy. As we navigate the COVID supply chain crisis, it’s worth considering how to reduce the severity of the next one.</p>]]></content:encoded>
    </item>
    <item>
      <title>The Case for White-Collar Apprenticeships</title>
      <link>https://alexkolchinski.com/2020/08/01/the-case-for-white-collar-apprenticeships/</link>
      <guid isPermaLink="true">https://alexkolchinski.com/2020/08/01/the-case-for-white-collar-apprenticeships/</guid>
      <pubDate>Sat, 01 Aug 2020 01:50:01 GMT</pubDate>
      <description>Over the past century, the labor market in America has seen a dramatic shift from blue-collar to white-collar work. According to the Bureau of Census Data, white-collar work in the US grew from 17.6% of total employment in 1900 to 59.9% in 2003 [1]. The “white-collar” categorization was then discontinued for lack of specificity, but […]</description>
      <content:encoded><![CDATA[<p class="wp-block-paragraph">Over the past century, the labor market in America has seen a dramatic shift from blue-collar to white-collar work. According to the Bureau of Census Data, white-collar work in the US grew from 17.6% of total employment in 1900 to 59.9% in 2003 [1]. The “white-collar” categorization was then discontinued for lack of specificity, but the fact remains that the majority of the American workforce is now employed in work where brainpower is more relevant than muscle power, a situation opposite to the way things were a century ago.</p>



<p class="wp-block-paragraph">As knowledge work was displacing manual labor as the most common form of employment in America, the amount of formal education completed by Americans also grew rapidly. Before World War 2, only a quarter of American adults had graduated high school, and a vanishingly small percent had gone to college. Americans like George Washington, Cornelius Vanderbilt and Thomas Edison rose to the heights of politics, business and science with little to no formal education. But that state of affairs changed quickly after the end of World War 2 with the introduction of the GI bill, which paid for postsecondary education for millions of veterans. 75 years later, a large majority of American adults now hold high school degrees, and over a third hold bachelor’s degrees.</p>



<figure class="wp-block-image"><img src="https://alexkolchinski.com/wp-content/uploads/2020/08/the-case-for-white-collar-apprenticeships-1.png" alt="" /></figure>



<p class="wp-block-paragraph"></p>



<p class="wp-block-paragraph">As Americans have begun to attain higher levels of formal educational credentials, so too have jobs begun to demand higher levels of credentials as a prerequisite. Some of this is because some jobs really do require extensive formal education: patients might not want to be operated on by a neurosurgeon who had learned exclusively on-the-job! But many white-collar jobs, like sales and clerical roles, which do not require extensive theoretical training and traditionally were not filled by college graduates, are starting to require higher degrees in a phenomenon often referred to as degree inflation [2]. </p>



<p class="wp-block-paragraph">Some of this is rational from the point of view of the employer, as college graduates can be expected to have higher levels of skills applicable to office work than high school graduates or dropouts, if only due to their four extra years of experience performing a type of knowledge work in college. But spending four years writing papers on art history with the goal of landing a sales job is not efficient from the point of view of either the prospective employer, who limits their labor pool and ends up paying a premium for entry-level labor, or the prospective employee, who incurs often-heavy tuition fees and four years of opportunity cost before starting a career.&nbsp;</p>



<p class="wp-block-paragraph">A better model exists in the form of apprenticeships. For centuries, teenagers have learned skilled blue-collar trades in collaboration with more experienced mentors, and have emerged into young adulthood as full-fledged professionals in their chosen field. As an added benefit, apprentices can be productive in the lower-skilled parts of a job almost from the get-go, implicitly paying for their own training and earning a wage as their counterparts in college instead pay tuition.</p>



<p class="wp-block-paragraph">This model has occasionally been applied to white-collar professions in countries like Germany and the UK, in which formal apprenticeships in the traditional mold exist for higher-skilled jobs like IT system administration and CNC machining. Those programs, where they exist, are generally coordinated to some degree by the relevant government, but America’s culture would likely pair better with a privatized apprenticeship model, which could fill niches unseen by a government administrator.</p>



<p class="wp-block-paragraph">In such a model, companies large enough to sustain internal apprenticeship programs would designate appropriate roles for which they would hire apprentices, like sales or web development. For those roles, they would hire recent high school graduates or even dropouts, who would commit to the apprenticeship program for several years.</p>



<p class="wp-block-paragraph">The apprentices would be assigned to a team and mentor, just like a typical intern or co-op student worker, with the key difference being that they would stay on each team for one to several years, rotating as appropriate to learn different aspects of their chosen profession.&nbsp;</p>



<p class="wp-block-paragraph">This on-the-job training would be paired with classroom training, where each cohort of apprentices would be instructed in relevant skills alongside their day job. This might include things like written and oral communication for sales apprentices, psychology and anthropology for marketers, and computer science and graphic design for web developers. Some of these courses could be conducted internally by the company’s more-senior employees; others could be outsourced to local colleges in a reverse co-op arrangement.</p>



<p class="wp-block-paragraph">The expectation would be that after completion of the apprenticeship program, some of its alumni would keep working in their new profession, while others would then go to college to pursue a broader and deeper formal education. Those that kept working would have the advantage of a four-year head start in their career trajectory and a much better financial situation than a recent college graduate; those that went to college after all would have a very compelling college admissions packet, transfer credits to the extent their employer could negotiate for them with universities, and a much better idea of what is worth studying in college than someone right out of high school. Talented students from poor families and poor schools might benefit to an especial degree from this sort of program, which would let them gain a stable financial footing and a better understanding than that provided by their high school of how to navigate a future college education. In this way, an apprenticeship program would play a similar role to that which the military plays for many teenagers today, but would prepare them for a rather different sort of work.</p>



<p class="wp-block-paragraph">Of course, some apprentices would probably flounder partway through the program, or conversely be tempted to leave for a competitor. Employers could protect their investment by a similar mechanism to what West Point does for cadets: apprentices could leave at any point during the first year with no strings attached, but those that left later on would have to pay back the expenses incurred in training them. The apprenticeship program could also have a contractual requirement to work for the employer for a certain number of years after completing the program, with the alternative option of paying a financial penalty. Employers could also protect their investments in their apprentices by paying the college tuition of apprenticeship alumni who wished to go to college afterwards, under the condition that they return for a certain number of years after completing their degrees – much like many employers pay for business school for their employees under the condition that they come back afterwards.</p>



<p class="wp-block-paragraph">This system of white-collar apprenticeships would have significant advantages for both the employer and the apprentice.&nbsp;</p>



<p class="wp-block-paragraph">The employer would be able to attract some of the most talented and driven teenagers with a unique value proposition and thereby gain a recruiting advantage over competitors that wait to hire much competed-over college graduates. The apprentices, once recruited, would also have the value of being able to perform necessary but less-skilled work that must currently be done by older employees for whom it is tremendously boring. Once done with their contractual term, a number of apprentices could be expected to stick around and keep working for the employer and delivering value for years to come, assuming the employer did a good enough job to keep offering opportunities for advancement and a good work culture.</p>



<p class="wp-block-paragraph">For high school students, the apprenticeship program would represent a unique opportunity to learn a profession with great career opportunities while earning a living straight out of high school, while keeping options open for a college education and even improving them. Right now, many high school graduates go off to college to study things which they will never use again, and make friends there with people who all too often end up moving to different cities and drifting apart over time. Instead, a well-run apprenticeship program would let students learn a gainful profession and the theoretical knowledge behind it while earning money from day one, and to use some of the most social years of their lives to form bonds with friends who would be far more likely to stay in the same industry and city and remain close social and professional contacts for decades. </p>



<p class="wp-block-paragraph">Starting a program like this would surely lead to howls of disapproval from those who see the one-size-fits-all track of formal education as the right way for everyone, but the company that started it would be more than compensated for that negative attention by the newfound stream of talent it would be able to access. Done right, an apprenticeship program would prove its worth in a matter of years, and would surely spawn numerous competitors – exactly what our economy needs in this era of degree inflation.</p>



<p class="wp-block-paragraph">The only question is, which company will seize the opportunity to go first? [3]</p>



<hr class="wp-block-separator" />



<p class="wp-block-paragraph">[1] <a href="https://www.encyclopedia.com/social-sciences/applied-and-social-sciences-magazines/employment-white-collar">https://www.encyclopedia.com/social-sciences/applied-and-social-sciences-magazines/employment-white-collar</a></p>



<p class="wp-block-paragraph">[2] <a href="https://web.archive.org/web/20231204195612/http://www.burning-glass.com/wp-content/uploads/Moving_the_Goalposts.pdf">https://www.burning-glass.com/wp-content/uploads/Moving_the_Goalposts.pdf</a></p>



<p class="wp-block-paragraph">[3] Thanks to Ani Mohan for <a href="https://blog.google/outreach-initiatives/grow-with-google/digital-jobs-program-help-americas-economic-recovery/">this link</a>: a recent Google press release mentions an apprenticeship program! Maybe my former employer will be the one to blaze this trail.</p>



<hr class="wp-block-separator" />



<p class="wp-block-paragraph"><em>Thanks to Allie Cavallaro, Ani Mohan, Alex Gruebele, Sal Calvo, Josh Pickering, and Anthony Buzzanco for helping edit drafts of this essay.</em></p>]]></content:encoded>
    </item>
    <item>
      <title>The Importance of India</title>
      <link>https://alexkolchinski.com/2020/06/22/the-importance-of-india/</link>
      <guid isPermaLink="true">https://alexkolchinski.com/2020/06/22/the-importance-of-india/</guid>
      <pubDate>Mon, 22 Jun 2020 15:51:38 GMT</pubDate>
      <description>This post sparked a great conversation on Hacker News! See the comments here. “Quantity has a quality of its own”–Attribution contested A significant factor in the power of states throughout history has been sheer numbers. Spain was able to control an overseas empire to a significantly greater degree than Portugal thanks in no small part […]</description>
      <content:encoded><![CDATA[<p class="wp-block-paragraph"><em>This post sparked a great conversation on Hacker News! See the comments <a href="https://news.ycombinator.com/item?id=23601597">here</a>.</em></p>



<p class="has-text-align-center wp-block-paragraph"><em>“Quantity has a quality of its own”</em><br>–Attribution contested</p>



<p class="wp-block-paragraph">A significant factor in the power of states throughout history has been sheer numbers. Spain was able to control an overseas empire to a significantly greater degree than Portugal thanks in no small part to its larger population. England subsequently rose in power and grew to control an empire of its own aided by a population boom in the 19th century. And the United States supplanted the United Kingdom as the world’s great power as its population vastly outgrew that of its former colonial overlord.</p>



<p class="wp-block-paragraph">Of course, many factors other than population influence the power of states, including economic productivity and the strength of institutions. But population is a multiplier for those factors, and a country with a large enough population can exercise comparable or greater power than more developed but less populous countries. For a long time, this was the role that Russia played in Europe: less advanced technologically and economically than much of Western Europe, but a great power by sheer force of numbers.&nbsp;</p>



<p class="wp-block-paragraph">The most important story in geopolitics today is the rise of China (population 1.3B) and its challenge to the supremacy of the United States (population 300M). The rise of China as a peer superpower to the United States is all but accomplished, and it appears that we are returning to a bipolar world, this time characterized by competition by the US and China, just as 1946-1991 was characterized by competition between the US and USSR.&nbsp;</p>



<p class="wp-block-paragraph">However, if China succeeds in sustaining its growth trajectory economically and militarily, it will grow to overshadow the United States, wielding the same ~5x population advantage over the US that the US now enjoys over its predecessor power, the United Kingdom. This also means that the repressive and authoritarian Chinese model will increasingly prevail over the free democratic model that America has championed.&nbsp;</p>



<p class="wp-block-paragraph">However, America is not the most populous democracy in the world. That honor belongs to India. India is forecast to surpass China in population in the next decade, and in the next few decades to grow almost 50% more populous than China. India currently punches below its weight on the world stage due to slow economic development: 30 years ago, its GDP per capita was similar to China’s, but is now 5x lower. However, if India were to enter a period of similarly high growth over the next 30 years as China has for the past 30, it would quickly become one of the most powerful countries in the world thanks to the scaling factor of its immense population. Moreover, India’s population is forecast to continue growing quickly, while China’s is forecast to shrink, and that will only compound any advantages that India accumulates.</p>



<p class="wp-block-paragraph">In a world that is quickly going from unipolar to multipolar, it is worth considering which states will wield influence in the century to come, and on behalf of which values (if any, other than self-interest!) they will wield it. If China rises to heights of power that eclipse the US completely, only India may be strong enough to speak for liberal democracy. It is therefore in the interest of the US and other Western powers to develop closer ties to India and encourage its economic development, so as to nurture a counterweight to the rise of authoritarianism in the 21st century.</p>



<p class="wp-block-paragraph"></p>]]></content:encoded>
    </item>
    <item>
      <title>What Can AI Really Do?</title>
      <link>https://alexkolchinski.com/2020/06/09/what-can-ai-really-do/</link>
      <guid isPermaLink="true">https://alexkolchinski.com/2020/06/09/what-can-ai-really-do/</guid>
      <pubDate>Tue, 09 Jun 2020 21:15:49 GMT</pubDate>
      <description>AI is creating tremendous change in the world, but it can be hard to tell where that’s truly happening and where there’s more hype than substance. This essay outlines how to tell one from the other.</description>
      <content:encoded><![CDATA[<h5 class="has-text-align-center wp-block-heading">Demystifying the magic of machine learning</h5>



<h6 class="has-text-align-center wp-block-heading">Alex Kolchinski – June 2020</h6>



<ul class="wp-block-list"><li><a href="#introduction">Introduction</a></li><li><a href="#history">The current wave of AI</a></li><li><a href="#background">Background</a></li><li><a href="#applications">Areas of AI research and their applications</a><ul><li><a href="#cv">Computer vision</a></li></ul><ul><li><a href="#audio">Audio</a></li></ul><ul><li><a href="#nlp">Natural language processing</a></li></ul><ul><li><a href="#rl">Reinforcement Learning</a></li></ul><ul><li><a href="#generativemodels">Generative models</a></li></ul><ul><li><a href="#recommendersystems">Recommender systems</a></li><li><a href="#optimization">Other optimization</a></li></ul></li><li><a href="#hype">The hype</a></li><li><a href="#examples">A few examples: real or hype?</a></li><li><a href="#agi">Artificial General Intelligence?</a></li><li><a href="#conclusion">Conclusion</a><ul><li><a href="#furtherreading">Further reading</a></li><li><a href="#acknowledgements">Acknowledgements</a></li></ul></li></ul>



<h2 class="wp-block-heading" id="introduction">Introduction</h2>



<p class="wp-block-paragraph">You’ve probably seen some pretty outrageous claims about what artificial intelligence (AI) can do. Maybe it’s been media headlines about how AI will soon be able to do everything humans can and will put us all out of work. Maybe it’s been company press releases about how their revolutionary AI technology will cure cancer or send rockets to the moon.</p>



<p class="wp-block-paragraph">Those claims, and many others like them, are largely or entirely untrue, and are designed to draw attention rather than to reflect what is actually possible or likely to be possible soon. But AI really has made incredible advances in the last decade, and will continue to drive tremendous changes across our economy and society in the years to come.&nbsp;</p>



<p class="wp-block-paragraph">Given that rapid progress, it’s important to be able to tell the outrageous claims about AI from those that are more credible, but how? Part of the reason I went to Stanford for grad school was to learn to do just that, and after doing quite a bit of AI research in different subfields, I’ve gained an intuition for what modern AI methods can really do. And by spending time in the world of startups, I’ve combined the more theoretical knowledge from the lab with a high-level understanding of how the advancements in research are being applied to the real world. I’ve found myself in many conversations with people wondering about exactly this question: which capabilities is AI quickly developing, and where will they be applicable? This is my attempt to distill those conversations into written form, and to spread a better understanding of the capabilities of AI in today’s world.&nbsp;</p>



<p class="wp-block-paragraph">I’ve written this essay to be approachable by those without a background in computer science or statistics, but you’ll probably get the most out of it if you have a quantitative background. You’ll probably find it especially interesting if you’re in a STEM field and are wondering how AI is likely to change the way things are done in your world, or are in business and entrepreneurship and are wondering which opportunities AI is now opening up, and which older ways of doing things it’s threatening to make obsolete!</p>



<p class="wp-block-paragraph"><strong>Update</strong>: Some readers have noted that this essay largely omits AI techniques other than deep learning. That&#8217;s true! Earlier waves of AI research are important both in a historical sense and because they&#8217;ve yielded many approaches that are still relevant to this day. Even after the advent of deep learning, many real-world problems are better solved by battle-tested techniques like logistic regression, decision trees, or any of a number of others. However, those techniques worked just as well ten years ago as they do today. The aim of this essay is to outline what is <em>newly</em> possible or becoming possible thanks to the current deep learning era of AI, and to draw a distinction between those real possibilities and problems that are either intractable or solvable without the need for deep learning techniques.</p>



<h2 class="wp-block-heading" id="history">The current wave of AI</h2>



<p class="wp-block-paragraph">In the decades following the emergence of computers, the term AI has shifted meaning repeatedly. Over and over, computers have gained new capabilities – like winning at chess, or answering simple questions – thanks to the application of new techniques, or just greater availability of processing power. When those capabilities appear novel enough, they often generate a wave of excitement around AI, and the public conversation around AI becomes centered on those new capabilities, largely to the exclusion of previously novel but now-mundane techniques that had generated previous waves of excitement.&nbsp;</p>



<figure class="wp-block-image is-resized"><img loading="lazy" src="https://alexkolchinski.com/wp-content/uploads/2020/06/what-can-ai-really-do-1.png" alt="" width="517" height="461" /><figcaption>A simple deep neural network. Source: Cburnett via Wikipedia, <a href="http://creativecommons.org/licenses/by-sa/3.0/">CC BY-SA 3.0</a> license</figcaption></figure>



<p class="wp-block-paragraph">The current wave of excitement centers on a set of techniques known as deep learning, which have unlocked unprecedented performance in a wide range of real-world applications. Deep learning rests on a surprisingly simple technique: if you stack many very simple functions (a couple of examples in one dimension are y = 2x, or y = tanh(x)) one after the other, you can nudge, or train, the resulting many-layered function, or neural network, to map complex inputs to outputs with surprising accuracy over time. For example, if you want to tell pictures of cats apart from pictures of dogs, you might train your network to output 1s for dogs and 0s for cats. To conduct one step of training, you might input a photo of a dog into the network, and then if the output was incorrectly closer to 0 than to 1, nudge the layer functions of the network in a direction that will make the final output of the network slightly smaller. For example, this might mean nudging a function that is currently y=2x to be y=1.9x. With some tuning, repeating this process thousands or even millions of times is likely to yield a network that can tell cats from dogs with surprisingly high accuracy.</p>



<p class="wp-block-paragraph">Of course, not all techniques currently referred to as AI are based on this deep learning approach, and the general excitement around AI has created plenty of nonsense — the joke goes that many startups tell the public that they’re an AI company, tell investors that they use machine learning, and internally use logistic regression (a simple and decades-old statistical technique) for some predictive task that may not even be core to their business model. A fictional example of this might be an online used-clothing marketplace, probably called something like <em>mylooks.ai</em>, that claims to be an AI company but in reality does nothing more than using simple statistics to predict when to send marketing emails to users to drive the most engagement.</p>



<div class="wp-block-image"><figure class="alignright is-resized"><img loading="lazy" src="https://alexkolchinski.com/wp-content/uploads/2020/06/what-can-ai-really-do-2.png" alt="" width="348" height="584" /></figure></div>



<p class="wp-block-paragraph">But don’t be fooled by the hype: deep learning really <em>has</em> changed the game in terms of what AI can do in the real world. Computational techniques before deep learning were very good at working with structured data (tables, databases, etc.), but much less good at unstructured data (images, video, audio, text, etc.), which is often very important in the real world. Deep learning, unlike older approaches, is very good at dealing with unstructured data, and that is where its power lies. Tasks that were previously hard or impossible to do reliably, like image identification, have suddenly become almost trivial, unlocking many applications. This comic (<a href="https://xkcd.com/1425/">https://xkcd.com/1425/</a>), from less than a decade ago, is already out of date: identifying a bird in a photo is now straightforward!</p>



<p class="wp-block-paragraph">All of these new abilities do come with a caveat: deep learning performance depends on huge amounts of data and computational power. The techniques of deep learning themselves have existed for decades, with limited applications like reading handwritten addresses for the postal service. But the limited availability of data and computational power hampered their performance until the 2010s, when two things happened. One was that the maturation of the Internet made vast amounts of text, images, and other unstructured data available. The other was the increasing performance of GPU chips, which were originally designed for gaming (I remember installing them in my gaming PC growing up!) but which, through a lucky accident, turned out to be incredibly useful for the acceleration of deep learning algorithms. When those two factors came together, deep learning made sudden and large gains in performance, which started drawing significant attention in 2012 when the AlexNet program smashed records on an image recognition challenge.</p>



<p class="wp-block-paragraph">The resulting attention drew in huge numbers of researchers and engineers in both academia and industry, and there has since been incredible progress in both fundamental AI research and downstream applications. Unfortunately, the attention has also created tremendous amounts of unwarranted hype, especially in industry but even in academia. The stakes are high to be able to tell one from the other, whether you’re an engineer deciding whether to work at a company that claims to be developing commercially-relevant AI, a policymaker forecasting changes in employment numbers, or an entrepreneur trying to tell a real opportunity from a mirage.&nbsp;</p>



<p class="wp-block-paragraph">So, what’s the best way to tell the real AI applications from the fake ones? The best strategy is to keep a finger on the pulse of what the research community is doing. With few exceptions, ideas that are deployed in industry are first published in the research literature, and then deployed (with modifications when needed) to a similar or analogous real-world task. Thus, an idea for an AI application that’s adjacent to something that’s already been published as a research finding is likely to be worth investigating, but one that’s totally unrelated to anything in the literature is best viewed with skepticism.&nbsp;</p>



<p class="wp-block-paragraph">There are significant nuances to which ideas are truly adjacent to each other – telling cats from dogs is an extremely similar task to telling defective machine parts from functional ones, but classifying a sentence as happy or sad is much easier than classifying it as insightful or inane. However, gaining a broad familiarity with what the different research communities in the world of AI are working on is a great way to gain an intuition for what is truly possible or likely to soon be possible with AI, and that’s what the rest of this essay is about.</p>



<h2 class="wp-block-heading" id="background">Background</h2>



<p class="wp-block-paragraph">Before we turn to a discussion of the communities in the world of AI research, it’s worth going over a few foundational ideas.&nbsp;</p>



<p class="wp-block-paragraph">A term that’s closely associated with AI is <strong>machine learning (ML)</strong>, which refers to AI approaches which learn from data, as opposed to being pre-programmed by humans. As most modern approaches to AI rely heavily or entirely on ML, the two terms have become synonymous in common usage.</p>



<p class="wp-block-paragraph">An AI <strong>model</strong> is a configuration of functions which can be, or have been, trained on some task. For example, AlexNet is an image recognition model composed of many functional layers. You might train a previously untrained copy of AlexNet on millions of photos to classify them into categories, or you might use a copy of AlexNet that’s already been trained on millions of images to help you tell cats from dogs.</p>



<p class="wp-block-paragraph">Deep learning techniques are applicable to a broad range of machine learning tasks, which can be roughly classified into many categories. A number of these, briefly described here, are commonly encountered, and worth knowing.</p>



<p class="wp-block-paragraph"><strong>Reinforcement learning</strong> (RL) involves step-by-step decision-making by a model, e.g. the controls software for a robot which plays ping-pong. At every time step, an RL model has some information about the state of the world (e.g. the position of the paddle and the ball) and takes some action (e.g. moving the paddle to the right) based on its policy, which is the term used to denote the program that picks actions based on states. The policy is trained to maximize a reward signal, which is encountered intermittently (e.g. +1 reward for winning a point, -1 reward for losing).</p>



<p class="wp-block-paragraph"><strong>Supervised learning</strong> is the setting in which a machine learning model is trained to map inputs to outputs. The model is trained with a labeled training set of known (input, output) pairs, and then tasked with predicting outputs corresponding to previously unseen inputs. Supervised learning can be further categorized into<em> </em><strong>classification</strong><em>, </em>where outputs are categories (“Is the animal in this photo a dog or a cat?”) and <strong>regression</strong>, where outputs are continuous (“How much does the dog in this photo weigh?”). <strong>Most current applications of deep learning in the real world fall under supervised learning.&nbsp;</strong></p>



<p class="wp-block-paragraph"><strong>Unsupervised learning</strong> is the setting in which there are no output labels in the training set, and the model’s task is instead to discover some structure in the input data. One common example of unsupervised learning is <strong>clustering</strong>, where the model is tasked with finding a grouping of the inputs (“Given these photos of animals, sort them into categories”).</p>



<p class="wp-block-paragraph">One reason why the distinction between supervised and unsupervised learning is important is data efficiency. Hiring humans to label enough data to train a machine learning model in the supervised setting can easily cost thousands of dollars, and so a number of techniques exist to improve data efficiency and label efficiency, including:</p>



<ul class="wp-block-list"><li><strong>Semi-supervised learning</strong>: Only some inputs in the training set have labels.</li><li><strong>Weakly supervised learning</strong>: Some labels in the training set are incorrect.</li><li><strong>Transfer learning</strong>: Use a model trained on data from task A to get higher performance on related task B, without needing to gather more data from task B. Often, the same more-general “task A” is commonly used to train more-specific “task B”s, e.g. an ImageNet image classifier, trained on millions of photos to identify hundreds of common objects, being fine-tuned on just a few thousand new images to learn a more specific task like telling cats from dogs.</li><li><strong>Self-supervised learning</strong> (now in vogue!): Learn the structure of a domain by training a model to learn a function where both the inputs and outputs are available from unlabeled training data, e.g. predicting missing words in a paragraph of text, using the other words as input, or predicting missing pixels in an image, using the other pixels as input. The resulting model can then be used for transfer learning, e.g. training a model on arbitrary photos in the self-supervised manner, then fine-tuning on relatively few cat and dog photos to tell them apart from each other. This is a very useful technique as in many domains, unlabeled data is plentifully available from sources like Wikipedia and YouTube.</li></ul>



<p class="wp-block-paragraph">Labels aside, the availability of the training data itself is the most important factor in the performance of deep learning on tasks which are amenable to it. For example, even if you know that deep learning techniques are great at telling images of objects apart from each other, and want to train a model to tell Martian rocks from Earth rocks, deep learning won’t do you any good unless you have a lot of photos of both! In academia, research into new deep learning methods is largely conducted on publicly available datasets, a constantly evolving set of which is in broad use by the community. Achieving state-of-the-art performance on one of those commonly used data sets is a sought-after achievement in academic research. In industry, on the other hand, ownership of a hard-to-gather dataset can be a key competitive advantage for a company whose technology is powered by AI, as the most effective deep learning algorithms are largely public knowledge but the right data to train them for a specific commercially valuable task can be very hard to source.</p>



<p class="wp-block-paragraph">The other key factor that dictates the performance of deep learning is the availability of computational power, often referred to as “compute” for short. Given a large enough training set of data, throwing more compute at the training process for a model will typically improve performance substantially. Indeed, achieving state-of-the-art (SOTA) results in some domains now costs hundreds of thousands of dollars in compute bills alone, and those numbers are only growing with time. A dynamic that this sometimes creates is that labs in industry, with their big budgets, will train huge models at great cost to achieve a SOTA result. Meanwhile, academic labs, with their more-modest resources, spend more of their effort on coming up with novel techniques that can drive higher and sometimes SOTA performance with less compute. A happy consequence of this second strain of research is that while achieving SOTA results keeps taking more and more compute and money, achieving the <em>same</em> level of performance on just about any task takes less and less compute with every passing year as the algorithms become more efficient.</p>



<h2 class="wp-block-heading" id="applications">Areas of AI research and their applications</h2>



<p class="wp-block-paragraph">Now that we’ve covered the broad categories and principles of machine learning, it’s time to dive into the most prominent areas of AI research. Application areas and types of machine learning intermingle freely: for example, ML for robotics may include both supervised learning for image recognition and reinforcement learning for controlling the robot.</p>



<p class="wp-block-paragraph">In practice, deep learning has unlocked huge gains in performance in some very specific areas, and knowing what these are is very useful for gauging which applications are likely to be fruitful. Something that is closely related to work in these areas is likely to be achievable with a bit of research and development (R&amp;D), while something totally unconnected is much more of a long shot in the near term.</p>



<h3 class="wp-block-heading" id="cv">Computer vision</h3>



<p class="wp-block-paragraph">Computer vision (CV) is the subfield of AI that deals with images, videos, and other related types of data like medical imaging. CV was the first field of AI to be revolutionized by the rise of deep learning, and it remains an extremely active area of research and applications. CV is also the most mature area of deep learning applications, and its high performance is well-understood and applicable to a number of tasks.</p>



<figure class="wp-block-image"><img src="https://alexkolchinski.com/wp-content/uploads/2020/06/what-can-ai-really-do-3.png" alt="" /><figcaption>The ImageNet Large Scale Visual Recognition Challenge. (Source: <a href="https://www.slideshare.net/xavigiro/image-classification-on-imagenet-d1l4-2017-upc-deep-learning-for-computer-vision/">Xavier Giro-o-Nieto</a>)</figcaption></figure>



<p class="wp-block-paragraph">Deep learning has been so successful in CV applications for a number of reasons. One is that visual data is unstructured, and deep learning is much better at handling unstructured data than earlier approaches. In addition, deep learning – with its hunger for data – has thrived thanks to the newfound plentitude of visual data. From images on Google Images to videos on YouTube, the Internet is now full of visual content sourced mostly from now-ubiquitous smartphones.&nbsp;</p>



<p class="wp-block-paragraph">Thanks to these factors, CV techniques powered by deep learning have been getting better and better at dealing with visual data. This is extremely important for downstream applications because visual data is able to capture a great deal of information about the physical world. Think of your own senses: while all are useful, sight is a particularly high-density way of gaining information about the world, from recognizing the objects around you to reading a book. Similarly, computers can now “see”, thanks to the advent of effective computer vision. This has unlocked all sorts of applications.</p>



<p class="wp-block-paragraph">The most straightforward of those application areas is in tasks based on image classification, which consists of sorting pictures into categories, e.g. cats vs. dogs. The research into image classification has advanced so far that computers often exceed human performance. This has unlocked all sorts of downstream applications: given the right data, you can identify people’s faces to grant them access to a building, identify the items in a retail customer’s shopping basket to charge them automatically as they leave the store without the need for a checkout lane, or automatically identify defective parts in a factory.</p>



<p class="wp-block-paragraph">Many more complex computer vision tasks exist as well. Two of common interest are object detection, which involves drawing a box around where certain objects are located in an image, and object segmentation, which involves precisely outlining the objects. Object detection has applications like identifying pedestrians in the field of view of an autonomous car’s camera(s). Object segmentation is useful for tasks like finding tumors in radiology images. If you want to locate objects in photos, that is now a very approachable problem.</p>



<figure class="wp-block-image"><img src="https://alexkolchinski.com/wp-content/uploads/2020/06/what-can-ai-really-do-4.png" alt="" /><figcaption>Object detection, from MTheiler via <a href="https://en.wikipedia.org/wiki/Object_detection">Wikipedia</a>. <a href="https://creativecommons.org/licenses/by-sa/4.0">CC BY-SA 4.0</a> license.</figcaption></figure>



<p class="wp-block-paragraph">Computer vision tasks are applicable to higher-dimensional data than 2D images as well. This includes things like 3D medical imaging and video. Video comes with its own set of challenges, including the need for huge amounts of compute due to the large number of frames. One common task in the world of CV for video is object tracking, or identifying where an object is from frame to frame of a video. This is useful in similar contexts as object detection in images, with the added capability of being able to handle movement; e.g. identifying a weed in the field of view of a <a href="https://www.youtube.com/watch?v=65_71XYsO64">weeding robot</a>, so as to precisely spray it with herbicide.</p>



<p class="wp-block-paragraph">Computer vision is rapidly gaining abilities on tasks which involve deeper inference as well. One of particular interest is pose estimation, where a CV model infers the positions of a person’s limbs and joints, or those of a robotic arm’s components, from a photo or video. This already works quite well and comes in handy for things like Snapchat filters, where being able to reliably identify the parts of someone’s face is important for then applying fun effects to it! Another intriguing research direction has been that of 3D reconstruction, or inferring the 3D shapes of objects from 2D images. There’s still plenty to be done there, but simpler “3D from 2D” inference for estimating sizes and distances to objects has been powering things like assisted braking in cars for over a decade.</p>



<p class="wp-block-paragraph">What’s particularly interesting about computer vision is that it can enable the use of commodity cameras and software (the computer vision model) in places where previously, more customized hardware or human labor would have been required. For example, a factory might install a camera at entryways that only admits workers whose faces are recognized as authorized employees and who are wearing an approved helmet. In this way, a camera and software could replace both ID card scanners and helmet checks. Many more such use cases have already been developed, and many more will be in the years to come.</p>



<p class="wp-block-paragraph">Computer vision can also serve as a surprisingly universal sensor for novel hardware, like <a href="https://www.gelsight.com/">enabling robots to “feel” objects</a> by visually measuring deformation in the membrane that’s in contact with the object in question.&nbsp;</p>



<p class="wp-block-paragraph">Of course, the fixed costs of training computer vision models for real-world tasks are usually quite high, due to the expense of both hiring researchers and engineers and collecting and labeling data (unless you are lucky enough to be able to use existing data like Wikipedia, but then your competitors will be too!) Deploying computer vision models, like deploying other machine learning models, also comes with the nontrivial variable costs of adjusting models to individual customers’ data and needs. These economics will dictate where computer vision is deployed in the next couple of decades, but expect to see it invisibly powering a wide range of applications across our economy in the decades to come.</p>



<h3 class="wp-block-heading" id="audio">Audio</h3>



<p class="wp-block-paragraph">The wide deployment of effective computer vision means that computers can now “The wide deployment of effective computer vision means that computers can now “see,” but they can also now “hear.” Just like visual data, audio data is complex and unstructured, which made it hard to work with before the rise of effective deep learning. And just like with visual data, deep learning has made it dramatically easier to work with audio – even more so than with visual data, as audio is simpler.</p>



<div class="wp-block-image"><figure class="alignright is-resized"><img loading="lazy" src="https://alexkolchinski.com/wp-content/uploads/2020/06/what-can-ai-really-do-5.png" alt="" width="251" height="445" /><figcaption>Via <a href="https://en.wikipedia.org/wiki/Siri#/media/File:Siri_on_iOS.png">Wikipedia</a></figcaption></figure></div>



<p class="wp-block-paragraph">This has powered the rise of a number of applications. Ten years ago, speech-to-text transcription was painfully inaccurate; now, Siri and Google Assistant are able to transcribe spoken commands and questions with surprisingly high accuracy. Working with music has changed completely as well. Ten years ago, Pandora suggested music to users based on an extensive database of hand-tagged information about songs. Now, Spotify combines that approach with algorithms that actually analyze the songs themselves with the help of deep learning to better match them to users’ tastes.</p>



<p class="wp-block-paragraph">Audio, while less studied in the research community than vision, is a very interesting field for AI applications because it’s the medium for human speech. Speech is in many ways the easiest and most natural way that we as humans communicate. That’s exactly why many tech companies are creating new platforms for audio interfaces, from Alexa speakers to Apple AirPods. Expect to see many more applications in the years to come, and to be interacting with computers more and more by talking to them. Numerous startups and big companies are working on the innovations to enable this change, and I expect many more to join them.</p>



<h3 class="wp-block-heading" id="nlp">Natural language processing</h3>



<p class="wp-block-paragraph">Alongside computer vision, the natural language processing (NLP) community is one of the most active in the world of deep learning. Broadly speaking, natural-language processing has to do with any task that primarily deals with human language, like when Siri answers questions or Google Translate translates text from one language to another.</p>



<p class="wp-block-paragraph">You may be wondering at this point why deep learning is applicable to natural language. After all, the types of unstructured data we’ve discussed so far are very different from language. Images, videos, and audio are all easy to represent in vector form – that is, as a long list of numbers. To simplify a bit, an image is a long list of pixel (dot) colors; a video is a long list of images, and an audio clip is a long list of sound intensities. But what about natural language? Each language is composed of some finite list of root words, and indeed, it’s possible to approach some NLP problems by assigning an integer index to each word and then training a relatively simple ML model not based on deep learning to solve the task, using the word indices as input data. This approach has worked well for some problems, like email spam filtering, but fails to capture the complexities that human language can express.</p>



<p class="wp-block-paragraph">But it turns out that there’s a way to represent words as vectors that allows more complex machine learning techniques, including deep learning, to work well with natural language and blow the performance of techniques that work directly with indexed words out of the water. That way is known as word embeddings. The principle is relatively simple: each word in a human language has some rate at which it co-occurs with every other word in the language, where a&nbsp; co-occurrence is when the two words are used with fewer than 5 (or some other small number) words between them, indicating some association between their meanings. For example, “apple” and “tree” have a high co-occurrence frequency in English, while “apple” and “tangent” have a much lower one. In this manner, it’s possible to pick a set of reference words – e.g. the 10,000 most common words in English, and a reference set of text – e.g. Wikipedia, and for each English word found in Wikipedia, count up the number of times it occurs within 5 words of each of the 10,000 reference words. This yields a 10,000-dimensional vector (list of numbers) of co-occurrence counts for every English word found in Wikipedia. The resulting list of word vectors can then be reduced to fewer than 10,000 dimensions – 100 is a common choice – without too much loss of information, in a manner similar to drawing a cube on a flat piece of paper. This then leaves us with a ~100 dimensional word vector for most words in the language.</p>



<figure class="wp-block-image"><img src="https://alexkolchinski.com/wp-content/uploads/2020/06/what-can-ai-really-do-6.png" alt="" /><figcaption>From <a href="https://nlp.stanford.edu/projects/glove/images/man_woman.jpg">Stanford NLP Group’s GloVe project</a></figcaption></figure>



<p class="wp-block-paragraph">These word vectors, also known as embeddings, have some very interesting properties. For one, words with similar meanings tend to cluster together in the 100 (or 50, or 200…) dimensional space they are embedded in such that “apple” and “orange” are closer together than “apple” and “knight”. This is because “apple” and “orange” are often found near the same words, like “tree” and “eat”, while “knight” is more likely to be found near “chess” and “horse”. Remarkably, analogies are often preserved in the n-dimensional space in which the word vectors live. For example, “king” and “queen” are likely to be separated by approximately the same distance, in the same direction, as “man” and “woman”. Finally, and most remarkably, it turns out that the shapes of the “clouds” of different languages’ word vectors are similar enough between languages that it’s possible to line them up and figure out quite accurately which word in e.g. French corresponds to “dog”, or “horse”, in English – just by lining up the clouds of all the words in English and French. In this way, it’s possible to translate between languages with some degree of accuracy with no examples whatsoever of translations between the languages! In the 19th century, it was necessary to find the Rosetta Stone to start deciphering Ancient Egyptian, but in 2020 it’s possible to decipher a hitherto unknown language using nothing more than a large body of text in that language alone – thanks to word embeddings.&nbsp;</p>



<p class="wp-block-paragraph">That said, word embeddings aren’t just useful in and of themselves; they also enable the application of deep learning to human language by mapping words to vectors, which deep learning techniques use as inputs. Ironically, in this case, mapping somewhat-structured data (individual words) to a less-structured but more numerical form (word vectors) enables higher performance on most NLP tasks.</p>



<p class="wp-block-paragraph">The range of these tasks is quite wide. One important NLP task is translation. While it’s possible to translate between languages using word vector cloud alignment as described above, or using the older methods that powered Google Translate for years, modern techniques based on deep learning have achieved much better levels of performance. Many other tasks are constantly being worked on by the research community, including question answering, which involves finding the answer to a question in a body of text, and sarcasm detection, which is exactly what it sounds like. A good sampler of tasks currently of interest to the research community is found in the <a href="https://w4ngatang.github.io/static/papers/superglue.pdf">SuperGLUE benchmark</a>, which is used to test the performance of NLP models. A broader list that gives a good overview of many NLP tasks is on this <a href="https://en.wikipedia.org/wiki/Natural_language_processing#Major_evaluations_and_tasks">Wikipedia page</a>.</p>



<p class="wp-block-paragraph"><strong>The capabilities of NLP techniques have seen incredible progress over the past two years, more so than any other field of machine learning.</strong> Much of this has been driven by the rise of effective transfer learning for NLP. Just like a computer vision model trained on a large number of photos of objects to classify those objects into categories can then be trained on a smaller number of photos of e.g. cats and dogs to tell cats from dogs, an NLP model trained on a general task can be fine-tuned to a more specific task as well. The more general task is typically some variation of predicting missing words in a piece of text. The more specific task could question answering, sarcasm detection, or any of a number of others.</p>



<p class="wp-block-paragraph">The trick is telling which tasks are well within the abilities of modern NLP, which tasks are on the horizon, and which are far in the future. Unfortunately, this is a difficult challenge, as human language itself is capable of expressing both very simple and very complex things.&nbsp;</p>



<p class="wp-block-paragraph">Simpler NLP tasks which rely on surface-level language features are now generally approachable with a high degree of accuracy. For example, <a href="https://www.grammarly.com/">Grammarly</a> has built a great business by making software that corrects word choice and sentence structure in a much more sophisticated way than traditional autocorrect – with technology powered by deep learning.&nbsp;</p>



<p class="wp-block-paragraph">However, tasks which rely on some understanding of the <em>meaning</em> of text are more complicated. The most important thing to remember when it comes to those tasks is that state-of-the-art NLP models only capture patterns of words, not deeper meanings. So, a good model will “know” that “Eiffel” and “Tower” are closely related, and even that “baguette” is likely to follow in a subsequent sentence. But it will be unable to tell whether the narrator is 10km, 100km, or 1000km from Lyon unless it has been trained on text specifically mentioning the distance from Paris to Lyon, as it does not have any true knowledge of the world. There is now quite a bit of research effort to address this problem and imbue NLP models with more explicit knowledge about the world, but these efforts have not yet changed the game.</p>



<p class="wp-block-paragraph">Despite this shortcoming, the state-of-the-art NLP models of 2020 are astonishingly powerful. The catch is that for complex tasks which require modeling the meaning of language, a model’s accuracy on that task will depend tremendously on the amount of specific data it is trained on. For example, a chatbot trained to answer common customer questions on a website might be able to achieve very high accuracy on a predefined and small set of interactions, if it is trained on something like 100,000 examples. But attempting to answer <em>all</em> customer questions would yield a much worse rate of appropriate responses, almost certainly low enough to be unacceptable for use by a business. For this and many other tasks, there is likely to be a “long tail” of examples that are especially hard for NLP models to handle. Leaving those to humans or some other fallback strategy while handling the more common examples automatically is a common strategy. This is what Amazon does with its customer support chat, when common queries are answered automatically but ones that the chatbot cannot answer confidently are routed to humans, and what Siri does by responding to only a small and fixed set of allowable questions and directing the rest to Google.</p>



<p class="wp-block-paragraph">Those sorts of approaches, with a strategy of biting off common but simple NLP tasks that currently consume significant human effort, or are cost-prohibitive to do with human labor entirely, and automating them while leaving the more complex tasks to humans, are starting to gain steam and are poised to have dramatic impact in industry the years to come. There are many examples of successful use cases already. Gmail automatically suggests completions for sentences and canned replies to simple emails, but lets the user choose what to include in their message. Legal software lets lawyers and paralegals search through documents and identify notable clauses with much more power and precision than was possible with keyword-only search in the recent past. And reams of documents in fields like transportation and logistics are starting to be processed automatically for later reference by humans, thanks to the application of computer vision and NLP to document processing.&nbsp;</p>



<p class="wp-block-paragraph">The recent advances made by NLP have both multiplied what it can do and reduced the difficulty of achieving high performance on many tasks. And the possible applications of these techniques are many. The channels that carry information between people are the sinews of the world economy, and people communicate in natural language, be that over the phone or electronic messaging. For the first time, we can work with that human-to-human information flow in scalable ways: reading an additional 100 articles about an industry per week used to mean hiring an additional analyst, and reading an additional 1000 used to mean hiring 10, but now, NLP summarization tools can summarize ten million articles almost as easily as they can summarize ten thousand.</p>



<h3 class="wp-block-heading" id="rl">Reinforcement Learning</h3>



<p class="wp-block-paragraph">Another extremely active area of research in the machine learning community, and one that has seen tremendous progress with the application of deep learning techniques, is the field of reinforcement learning (RL). Reinforcement learning is unlike other paradigms of machine learning in that it involves step-by-step decision making towards some goal. Some example tasks where an RL framework is relevant include playing video games and board games, and controlling robots and autonomous vehicles. In each case, there is a changeable state (e.g. the positions of pieces on a chessboard, or the pixels of an Atari game’s screen), possible actions that can be taken (legal moves in chess, or the steering input to an autonomous vehicle), and a “reward” that denotes a successful or unsuccessful outcome (e.g. a positive reward for winning a chess game, a negative reward for crashing a car) which is used as the signal to update the model.&nbsp;</p>



<p class="wp-block-paragraph">The reward in RL is what frames the training process for the model. In the same way that a supervised learning model’s underlying functions are “nudged” when it makes mistakes on training examples to reduce similar mistakes in the future, an RL model’s underlying functions are “nudged” to be more likely to take actions that led to positive rewards, and less likely to take actions that led to negative rewards.</p>



<figure class="wp-block-image"><img src="https://alexkolchinski.com/wp-content/uploads/2020/06/what-can-ai-really-do-7.png" alt="" /><figcaption>Kaspar, M., Osorio, J. D. M., &amp; Bock, J. (2020). Sim2Real Transfer for Reinforcement Learning without Dynamics Randomization. <em>arXiv preprint arXiv:2002.11635</em>.</figcaption></figure>



<p class="wp-block-paragraph">Reinforcement learning has seen tremendous advances in recent years thanks to the heavy application of deep learning techniques, with modern approaches (deep RL) able to <a href="https://www.youtube.com/watch?v=g-dKXOlsf98">defeat humans at the board game <em>Go</em></a> (previously the last stronghold of human superiority over computers in popular board games) and score highly at many video games.&nbsp;</p>



<p class="wp-block-paragraph">However, as with other domains, the performance of deep learning techniques in RL is chiefly limited by the availability of data. And more so than in other domains, collecting data for training RL algorithms is <a href="https://www.youtube.com/watch?v=iaF43Ze1oeI&amp;feature=emb_title">fiendishly difficult</a>, as the step-by-step nature of the setting often requires millions of iterations of alternatively applying the model to its target task and retraining it on the data gathered from those applications.</p>



<p class="wp-block-paragraph">This means that in settings where running RL models and collecting data on their performance is cheap and fast – i.e. in settings that can exist entirely in a computer, like video games and board games – deep RL algorithms have shown tremendous gains in performance. But in settings where interaction with the real world is necessary for measuring the performance of models and collecting the data with which to improve their performance – like robotics and online education – the rate at which data can be collected is dramatically lower than in the fully virtual settings, and this has crippled the performance of deep RL in those settings. As a result, in settings like robotic control and autonomous vehicles, controls software is still not commonly powered by deep learning despite the power of those techniques in other areas.</p>



<p class="wp-block-paragraph">There is now a significant push to close the gap between RL in virtual environments and RL in the real world, with one interesting direction being the use of simulation and its mapping to the real world, otherwise known as sim2real. For example, a deep RL algorithm for robotic control might be trained on a huge number of runs on a simulated robot, then fine-tuned on the real physical robot for thousands rather than millions of runs. These techniques show promise, but are still far from perfect. For example, OpenAI <a href="https://www.youtube.com/watch?v=jm-ihc7CASY">recently demonstrated</a> a robotic hand that had been trained with a sim2real approach to solve a Rubik’s cube. Training the hand to work perfectly in simulation took them about 2 months, but getting it to then work in the real world took another 2 years, and the resulting performance was still extremely unreliable and only worked for configurations of the cube that required few moves to solve.</p>



<p class="wp-block-paragraph">Another approach which is useful in addressing the data-inefficiency of reinforcement learning is known as imitation learning. Models trained by reinforcement learning typically blunder about for a long time after training starts, learning haphazardly to take actions that lead to positive rewards. Imitation learning instead trains models to choose actions that roughly imitate demonstrations of the task at hand, which are typically sourced from humans. One approach is to train a model with imitation learning on some number of human demonstrations (e.g. of people playing Tetris) and then to switch to trial-and-error reinforcement learning to further improve performance. This is apparently the approach taken by Pieter Abeel’s startup <a href="https://covariant.ai/">Covariant AI</a>, which is attempting to train robot arms to move objects between bins in warehouses – a very hot area right now.</p>



<p class="wp-block-paragraph">The bet that Covariant and its many competitors are placing is that while “RL for the real world” is not working well enough for most applications yet, it is on the cusp of a breakthrough that will unlock many new applications and billions of dollars of value. Indeed, RL is showing new and impressive results on a regular basis, from the abovementioned Rubik’s Cube robot hand to Google’s <a href="https://ai.googleblog.com/2020/04/chip-design-with-deep-reinforcement.html">recent result</a> showing success in designing layouts for computer chips (a very complex task) with the help of RL.&nbsp;</p>



<p class="wp-block-paragraph">When the performance of RL crosses into industrially-useful territory, it is likely to follow a similar pattern as NLP technology: deployment in partnership with humans. Just as a chatbot might be able to handle 50% of customer inquiries automatically and route the rest to human agents, a bin-picking robot for warehouses might be able to handle 99% of objects and require humans walking around the warehouse to handle the 1% that are the hardest to pick up.&nbsp;</p>



<p class="wp-block-paragraph">It is still unclear whether RL will cross the line into widely-useful performance in one year or in ten, but when it does, expect many consequences. Not the least will be much more adaptable robots, which will mean a number of use cases outside the factories and warehouses where most robots are currently deployed. Roombas may not be the only robots people encounter in their day-to-day lives for long!</p>



<h3 class="wp-block-heading" id="generativemodels">Generative models</h3>



<figure class="wp-block-image"><img src="https://alexkolchinski.com/wp-content/uploads/2020/06/what-can-ai-really-do-8.png" alt="" /><figcaption>From a 2017 model (WGAN)&#8230; &#8211; from<a href="https://github.com/shekkizh/WassersteinGAN.tensorflow"> shekkizh on GitHub</a>, MIT License</figcaption></figure>



<p class="wp-block-paragraph">So far, we’ve discussed many ways in which deep learning models are able to process unstructured input data, whether it’s in the form of images, audio, video, text, or something else entirely. However, deep learning has enabled something else that’s quite remarkable: training models that <strong>output </strong>unstructured data.&nbsp;</p>



<p class="wp-block-paragraph">What this means is that it’s now possible to train a model to generate new images, audio, text, etc. This can take the form of training a model on many images of people’s faces to generate new images of people who don’t exist at all. It can also take the form of training a model on Wikipedia articles to output plausible-seeming sentences and paragraphs. This works both in the non-conditional setting (“Given these millions of photos of people’s faces, come up with more that look similar”) and in the conditional setting (“Given these millions of paragraphs with accompanying audio clips of people reading them, learn to take a new paragraph and generate audio that sounds like someone reading it out loud”).</p>



<p class="wp-block-paragraph">Just like NLP, generative models have seen tremendous advances in the last couple of years, and are now showing incredible results. It’s now possible to do things like generate extremely realistic images from nothing more than a text description <a href="http://nvidia-research-mingyuliu.com/gaugan/">or a sketch</a>, or a video of a news anchor or politician giving a speech that they never gave at all (the controversial “deep fakes”), or to automatically generate captions for pictures and videos.</p>



<figure class="wp-block-image"><img src="https://alexkolchinski.com/wp-content/uploads/2020/06/what-can-ai-really-do-9.png" alt="" /><figcaption>To a 2018 model! More recent models have gotten even better. From Karras, T., Aila, T., Laine, S., &amp; Lehtinen, J. (2017). Progressive growing of gans for improved quality, stability, and variation. <em>arXiv preprint arXiv:1710.10196</em>.</figcaption></figure>



<p class="wp-block-paragraph">The potential for downstream applications are many. One area that’s already seeing substantial application of generative models is computational photography. It used to be the case that cell phone cameras were much less capable than full-fledged DSLRs with large lenses. But it turns out that in many cases, the software on the phone can more than correct for the shortcomings of the hardware. Deep learning like image denoising and image superresolution are able to dramatically increase the perceived quality of photos by inferring missing details, and companies like Google and Apple have already deployed these or similar techniques to their phones to enable things like professional-quality portraits and incredible low-light photography. One funny consequence may be the reliability of details in photos: for example, as more sophisticated algorithms are deployed to phones, a photo taken at dusk may look as good as one taken in the daytime, but the colors of objects may be slightly or even significantly wrong as the algorithm “guesses” those that there is not enough light to “see” properly.&nbsp;</p>



<p class="wp-block-paragraph">Other artistic applications of generative models are sure to follow. Creative tools are likely to be some of the first areas to be disrupted. Imagine an alternative to PhotoShop where instead of having to manually airbrush out a pole and fill in the background, there was a button to “remove object” in one click that did it automatically. Or imagine an animation program that could automatically design a scene based on text instructions and place characters in it, then have those characters perform actions and express emotions based on high-level commands – removing the need to manually design much at all. Or a music production program that could automatically adjust the style of a song to be more exciting, or more menacing, or add incredibly realistic instruments to accompany a singer dynamically. These tools are likely to start out with much less room for customization and lower-quality output than existing high-effort tools, but enable hobbyists (even children and teenagers!) to produce songs and movies that would have previously taken much more skill and labor. This could have as much of an effect on the landscape of content production as the advent of the Internet and YouTube, SoundCloud, etc. Eventually, I expect the classic disruptive technology cycle to take place and the high-leverage creative tools to make their way upmarket and take over more and more of the creative software space.</p>



<figure class="wp-block-image"><img src="https://alexkolchinski.com/wp-content/uploads/2020/06/what-can-ai-really-do-10.png" alt="" /><figcaption>The 12 photos in the bottom right were generated from the sketches in the top row to imitate the styles of the photos in the leftmost column! From Park, T., Liu, M. Y., Wang, T. C., &amp; Zhu, J. Y. (2019). Semantic image synthesis with spatially-adaptive normalization. In <em>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</em> (pp. 2337-2346).</figcaption></figure>



<h3 class="wp-block-heading" id="recommendersystems">Recommender systems</h3>



<p class="wp-block-paragraph">Recommender systems may well be the AI application that has already driven the most economic impact. Their purpose is to match content with users, whether that content is a video on YouTube, a movie on Netflix, or an ad on Google or Facebook.&nbsp;&nbsp;</p>



<p class="wp-block-paragraph">Recommender systems are a key part of the “secret sauce” of a number of large companies, and often rely on proprietary data, so their development has mostly taken place in secret within the research labs of large companies, with little activity in academia. Thus, it’s hard for outsiders to tell just how capable they are, and what techniques are being used to advance the state of the art in the field. However, it’s safe to say from the incredible amounts of money made by Google, Facebook, Netflix and others that the systems work very well. It’s also a safe bet that deep learning has been having an impact on the field, whether by automatically analyzing the content of YouTube videos or even by enhancing the matching algorithms themselves.</p>



<h3 class="wp-block-heading" id="optimization">Other optimization</h3>



<p class="wp-block-paragraph">Recommender systems are not the only “less sexy” under-the-radar place where modern AI can generate significant economic impact. Sophisticated quantitative finance firms like Renaissance and D. E. Shaw have likely been using deep learning to gain a competitive edge in the markets for much longer than those techniques have been in the public eye — likely since the 90s or even earlier.</p>



<p class="wp-block-paragraph">But many other parts of the economy also rely on solving optimization problems to make money, from FedEx package routing to the operation of electrical grids. The optimization of these business problems, often referred to as operations research, has been the topic of study since at least the 1940s, and has met with great success. In many cases, well-established approaches from computer science and statistics, like linear programming, are more than adequate to find near-optimal solutions to business optimization problems.</p>



<p class="wp-block-paragraph">But there are also cases where it will be possible to find better solutions to business optimization problems by finding complex patterns in high-dimensional data: exactly where deep learning excels.&nbsp; And even in cases where applying deep learning techniques can yield only a small edge over more traditional techniques, each percentage point of increased efficiency can be worth billions of dollars. As such, I expect that a number of companies are already hard at work in this space and that more will emerge in the years to come.</p>



<h2 class="wp-block-heading" id="hype">The hype</h2>



<p class="wp-block-paragraph">Alongside the places where deep learning-powered AI is truly driving huge progress, there are a number of realms where perception of the capabilities of AI far surpasses what it can actually do. Much of this is self-reinforcing: the more that everyone believes that AI is going to acquire incredible powers, the more they pay attention to it, and this creates incentives for the media to hype up the power of AI to draw clicks, and for businesspeople to hype up their usage of AI to draw media attention and investors dollars. Even AI researchers are often tempted to overstate the generality of their results to generate enthusiasm and win grants.</p>



<p class="wp-block-paragraph">Interestingly, this cycle has played out a number of times before. Each time, a wave of progress in the genuine capabilities of AI excited public sentiment, which in turn incited a wave of overly enthusiastic promises from the media, business, and research communities. Eventually, expectations rose so high that after the technology inevitably failed to fulfil most of them, an “AI winter” followed, with enthusiasm and funding depressed for years.</p>



<p class="wp-block-paragraph">I expect a similar dynamic to play out soon in public sentiment after AI fails to deliver most of the more outrageous things that have been promised over the past few years. However, I don’t expect a full “winter” to take place, as AI techniques are now creating incredible economic value as detailed in this essay, and that stream of money will continue to stimulate interest in AI in the business and research communities. Before long, we should reach a point where expectations of AI’s abilities are better matched to reality.</p>



<p class="wp-block-paragraph">In the meantime, it’s best to treat the bolder claims around AI with a skeptical eye. Again, the best way to tell if a claim about a new AI application has a realistic chance of being true is to consider whether it lines up well with an existing research direction, or is similar enough that existing techniques that perform well are likely to generalize to it.</p>



<h2 class="wp-block-heading" id="examples">A few examples: real or hype?</h2>



<p class="wp-block-paragraph">Now, perhaps a few example claims are in order! Here, I’ll outline my opinions on a few that I’ve heard, either exactly or approximately, over the past couple of years. This list is of course far from exhaustive, but I hope it’s a useful sample of more and less realistic claims about what AI can do.</p>



<p class="wp-block-paragraph">“<strong>We’re going to build an AI that can defeat humans at most video games.”</strong>: This one seems likely! RL techniques are well-suited to learning high performance on video games, and that performance is rising quickly. Of course, AI currently performs much better on some games than others, and there are likely to be games where humans have an edge for the foreseeable future.</p>



<p class="wp-block-paragraph"><strong>“We’re going to build an AI that will be able to drive anywhere a human can.” </strong>This claim is one where reality has fallen far short of expectations. Training AI to interact with the physical world turns out to be a hard and messy problem, and driving human passengers is an area where mistakes are particularly costly. Expect to see autonomous driving roll out only in limited ways in the near future &#8211; e.g. small, slow food delivery vehicles which are unlikely to hurt anyone in a crash and can be remote-controlled in the case of unexpected circumstances (another example of AI deployment in partnership with humans!); Nuro is working on something like this. Highway trucking in good weather is another task that’s easier than full autonomy, and will have tremendous economic impact when even partially automated.</p>



<p class="wp-block-paragraph">“<strong>We’re going to build an AI that will find drugs to cure diseases.”</strong> In theory, finding small molecules or biologics to bind to a known target to affect the course of a disease is a problem that can be solved computationally. In practice, computational approaches to drug discovery are at best an aid to wet-lab work. Could that change? Absolutely, but I haven’t yet heard of any applications of deep learning that are redefining performance in drug discovery in the same way that they did in image classification. Gathering data in this space is hard, simulations of binding affinity are far from perfect, and there is so much known structure in the domain – the laws of physics and chemistry – that deep learning, which is particularly good at discovering hard-to-find structure in data, may not be the right tool for the job. Computational modeling of drug candidates will improve, possibly with some help from deep learning techniques, but it is unlikely to replace <em>in vitro</em> and <em>in vivo</em> testing entirely due to the complexity of biological systems. This will change only when we can simulate a human body down to the molecular level, something unlikely to happen for many decades.<br><strong>Update:</strong> Jonas Köhler has contributed a thoughtful <a href="https://www.reddit.com/r/MachineLearning/comments/gzwncp/d_what_can_ai_really_do/">post</a> on Reddit disagreeing with some of my arguments here; I highly recommend taking a look!</p>



<p class="wp-block-paragraph"><strong>“We’re going to build an AI that will teach students.”</strong> This one sounds great, and my hope of making progress in this direction is what led me to start a PhD at Stanford. However, it turns out that education is an incredibly complex and social task, and one that it is tightly coupled with just about everything that makes us human. We as humans have evolved to be able to infer each other’s thoughts and feelings in ways that let us teach each other effectively as well. This, however, is far from the capabilities of AI right now: if even realistic chatbots are out of reach right now, what hope is there for an electronic Socrates? There is educational software now that is reaching millions of learners (think Duolingo and Anki), and adaptive algorithms do contribute to its performance. However, I haven’t seen any cases so far where deep learning techniques outperform simple adaptive strategies like moving students to a new topic after they get three questions in a row right on the previous one. It may be that with more and more students interacting with educational software, the vast volume of data will yield significantly improved algorithms, but that’s going to be a tough nut to crack at best.</p>



<h2 class="wp-block-heading" id="agi">Artificial General Intelligence?</h2>



<figure class="wp-block-image"><img src="https://alexkolchinski.com/wp-content/uploads/2020/06/what-can-ai-really-do-11.png" alt="" /></figure>



<p class="wp-block-paragraph"><br>Finally, I would be remiss if I didn’t mention the debate around artificial general intelligence (AGI), also known as strong AI. AGI is essentially AI that can do everything a human can – and probably much more. Every time the state of AI advances, there tends to be alarm about the impending rise of AGI, like the famous example in the Terminator movies when the Skynet AI gains general intelligence and takes over the world. But in reality, we are almost certainly far from the advent of AGI. Human intelligence is still poorly understood, and so simulating it in a computer is not something that’s approachable directly.&nbsp;</p>



<p class="wp-block-paragraph">It’s true that individual functions of the human brain, like visual and auditory perception, are now well-approximated by deep learning methods. But other faculties, like abstract reasoning, are still all-but-unapproachable by any ML techniques. Could this change, just as vision went from being unapproachable to almost trivial? Certainly, but that would require at least one and probably a number of breakthroughs in ML. And even if we learn to approximate all the functions of the human brain individually with ML, it will be some time before we stitch those programs together into a coherent whole and learn to train that model in a way that will teach it as much about the world as humans learn over a lifespan of decades. If AGI is gradually cobbled together from different advances in ML as it is now practiced, we will have plenty of warning, and the emergence is likely to be gradual enough that there is no single identifiable “before” and “after.”</p>



<p class="wp-block-paragraph">It’s also possible to sidestep the complexities of duplicating the functions of human cognition one-by-one with ML and try to directly simulate a human brain neuron-by-neuron, but this too is a daunting problem. A brain is an incredibly complex system, not yet fully understood, and researchers have thus far been unable to simulate the nervous system of any creature more complex than a worm. There will surely be progress here as well, but here too it will be incremental. If there comes a day years or more likely decades from now when we build a supercomputer that successfully embodies a conscious human mind by simulating its physical structure, we will have hints that such a thing is in the works well before it actually happens, just like we would if existing approaches to ML gradually add up to full AGI.</p>



<p class="wp-block-paragraph">Could we stumble into AGI in some way that doesn’t require decades of progress? The only way I can imagine this happening is if current deep learning methods turn out to be far more powerful than imagined. This isn’t entirely impossible: current methods have shown a remarkable ability to gain performance in areas like NLP from nothing more than constant increases in model size and computational power and data available for training the model. The company OpenAI claims to be pursuing exactly this strategy, apparently with a focus on reinforcement learning and similar approaches – the idea seems to be to expose increasingly huge RL models to increasingly complex real-world problems and see if AGI emerges.</p>



<p class="wp-block-paragraph">However, it seems unlikely that these constant increases in power applied to existing techniques will yield a sudden breakthrough in performance that unlocks AGI. RL algorithms are so far incapable of so much as reliably solving a physical Rubik’s Cube. There seems to be no reason to believe that they will suddenly learn to do everything humans can. It is of course possible that this happens, and with that eventuality in mind it’s worth thinking through how best to manage the rise of AGI when it comes about – it will indeed be an incredibly powerful technology, maybe the most powerful ever in human history. But the day it comes into being is probably still decades away, if not longer.</p>



<h2 class="wp-block-heading" id="conclusion">Conclusion</h2>



<p class="wp-block-paragraph">In this time of quick progress in the capabilities of AI, it’s especially important to be judicious in telling new capabilities of the technology – and the opportunities they open up – with marketing hype, designed more to attract attention than to represent reality. Don’t be fooled by overly optimistic narratives, but also don’t assume that AI is all snake oil: there are many frontiers opening up ahead of us. For my fellow researchers and businesspeople reading this who will be opening those frontiers in the future, I wish you luck!</p>



<p class="wp-block-paragraph"><strong>If you have any comments, please leave them below! I also welcome any messages at (My first name) @ (My last name).com – get in touch if you want to chat about AI, startups, or anything else that’s on your mind.</strong></p>



<h4 class="wp-block-heading" id="furtherreading">Further reading</h4>



<p class="wp-block-paragraph">There are vast quantities of writing on AI on the Internet today, but only some of it is of high quality. If you’re looking for more places to read about a high-level view of AI, I recommend reading around the <a href="https://a16z.com/category/machine-learning/">Andereesen Horowitz site and blog</a>, <a href="https://www.ben-evans.com/">Benedict Evans’s newsletter</a>, and <a href="http://karpathy.github.io/">Andrej Karpathy’s blog</a>.</p>



<h4 class="wp-block-heading" id="acknowledgements">Acknowledgements</h4>



<p class="wp-block-paragraph">Many thanks to Anthony Buzzanco, Allie Cavallaro, Alex Gruebele and Anna Kolchinski for reading and editing drafts of this essay. This would have been far less readable without their help! </p>]]></content:encoded>
    </item>
    <item>
      <title>A bird’s-eye view of modern AI from NeurIPS 2019</title>
      <link>https://alexkolchinski.com/2019/12/30/neurips-2019/</link>
      <guid isPermaLink="true">https://alexkolchinski.com/2019/12/30/neurips-2019/</guid>
      <pubDate>Mon, 30 Dec 2019 04:21:33 GMT</pubDate>
      <description>Table of Contents Introduction Robustness and generalizability Efficiency Data efficiency Computational efficiency Directions and applications Graph neural networks Reinforcement learning and contextual bandits Natural language processing SysML Generative models Miscellanea Conclusion Introduction This year, I had a chance to attend NeurIPS, the most prominent conference in artificial intelligence and machine learning (AI/ML), to present a […]</description>
      <content:encoded><![CDATA[<h5 class="wp-block-heading"><strong>Table of Contents </strong></h5>



<ul class="wp-block-list"><li><a href="#introduction">Introduction</a></li><li><a href="#robustness">Robustness and generalizability</a></li><li><a href="#efficiency">Efficiency</a><ul><li><a href="#data-efficiency">Data efficiency</a></li><li><a href="#compute-efficiency">Computational efficiency</a></li></ul></li><li><a href="#directions">Directions and applications</a><ul><li><a href="#graph-nns">Graph neural networks</a></li><li><a href="#rl">Reinforcement learning and contextual bandits</a></li><li><a href="#nlp">Natural language processing</a></li><li><a href="#sysml">SysML</a></li><li><a href="#generative">Generative models</a></li><li><a href="#misc">Miscellanea</a></li></ul></li><li><a href="#conclusion">Conclusion</a></li></ul>



<h2 class="wp-block-heading" id="introduction">Introduction</h2>



<p class="wp-block-paragraph">This year, I had a chance to attend NeurIPS, the most prominent conference in artificial intelligence and machine learning (AI/ML), to present a workshop paper. I’ve spent the past couple of years working on a combination of AI research in various subfields and tech startups and so have been following the evolution of AI with interest. This conference, bringing together as it does some of the best researchers and practitioners in the field, was a particularly good vantage point to gauge the state of, and changes in, how people are thinking about and using AI. Here, I’ve collected some of my impressions in the hopes that they might be useful to others. If you&#8217;re curious about other people&#8217;s perspectives, Andrey Kurenkov collected some links to <a href="https://david-abel.github.io/notes/neurips_2019.pdf">various</a> <a href="https://github.com/RobertTLange/conference-school-notes/tree/master/2019-12-NeuRIPS">talks</a> <a href="https://medium.com/@natolambert/reflections-on-neurips-2019-6317f102ee09">and</a> <a href="https://huyenchip.com/2019/12/18/key-trends-neurips-2019.html">key trends</a> in his <a href="https://thegradient.pub/neurips-2019-too-big/">recent post</a>, which is also worth a look.</p>



<p class="wp-block-paragraph">The most overarching theme I noticed at NeurIPS was the maturation of deep learning as a set of techniques. Since AlexNet won the ImageNet challenge resoundingly in 2012 by applying deep learning to a contest previously dominated by classical computer vision, deep learning has attracted a very large share of the attention within the field of AI/ML. Since then, the efforts of countless researchers developing deep learning and applying it to various problems have accomplished things like beating humans at Go, training robotic hands to solve Rubik’s cubes, and transcribing speech with unprecedented accuracy. Successes like these have generated excitement both within the AI community and elsewhere, with the mainstream impression tending towards an overestimate of what AI can actually do, fueled by the more narrowly circumscribed successes of new, largely deep-learning powered, methods. (Gary Marcus has a great recent <a href="https://thegradient.pub/an-epidemic-of-ai-misinformation/">essay</a> talking about this in more detail.)</p>



<p class="wp-block-paragraph">However, a perspective that I find more useful than “the robots are coming” is the one I heard from Michael I. Jordan when he came to Stanford to give a talk in which he described modern machine learning as the emerging field of engineering which deals with data. Consistent with this perspective, I saw a number of lines of inquiry at NeurIPS which are developing the field into more nuanced directions than “Got a prediction problem? Throw a deep net at it.” I’ll break down my impressions into three general areas: making models more robust and generalizable for the real world, making models more efficient, and interesting and emerging applications. While I don’t claim that my impressions are a representative sample of the field as a whole, I hope they will prove useful nonetheless.</p>



<h2 class="wp-block-heading" id="robustness">Robustness and generalizability</h2>



<p class="wp-block-paragraph">One prominent category of work that I saw at NeurIPS was that which addressed real-word requirements for successfully deploying models other than just high test-set accuracy. While a canonical case of a successful deep learning model, like an image classifier trained on the ImageNet dataset, is successful within its own domain, the real world in which models must be grounded and deployed is complex and ambiguous in ways which models much address if they are to be useful in practice.</p>



<p class="wp-block-paragraph">One of these complexities is calibration: the ability of a model to estimate the confidence with which it makes predictions. For many real-world tasks, it’s necessary not only to have an argmax prediction, but to know how likely that prediction is to be accurate, so as to inform the weight given to that prediction in subsequent decision-making. A number of papers at NeurIPS addressed better approaches to this complexity.</p>



<p class="wp-block-paragraph">Another complexity is ensuring that models are assigning appropriate importance to features which are semantically meaningful and generalizable, which in one way or another includes representation learning, interpretability and adversarial examples. A story I heard that illustrates the motivation for this line of research had its origins in a hospital, which had created a dataset of (if I remember correctly) chest X-ray images with associated labels of which patients had pneumonia and which did not. When researchers trained a model to predict the pneumonia labels, its out-of-sample performance was excellent. However, further digging revealed that in that hospital, patients likely to have pneumonia were sent to the “high-priority” X-ray machine, and lower-priority patients were sent to another machine entirely. It also emerged that the machines left characteristic visual signatures on the scans they generated and that the model had learned to use those signatures as the primary feature for its predictions, leading to predictions that were not based off of anything semantically relevant to pneumonia status and which would neither yield incremental useful information in the original hospital nor generalize in any way to other hospitals and machines.&nbsp;</p>



<p class="wp-block-paragraph">This story is an example of a “clever Hans” moment, in which a model “cheats” by finding a quirk of the dataset it is trained on without learning anything meaningful and generalizable about the underlying task. I had a great conversation about this with Klaus-Robert Müller, <a href="https://www.nature.com/articles/s41467-019-08987-4">whose paper</a> on the phenomenon is well worth a read. I saw a number of other papers at NeurIPS dealing with interpretability of models, as well as representation learning, the related study of how models represent data. A notable subset of this work was in disentangled representations, an approach which aims to induce models to learn representations of data which are composed of meaningfully and/or usefully factorized components. An example would be a generative model of human faces which learns latent dimensions corresponding to hair color, emotion, etc., thus allowing better interpretability and control of the task.</p>



<p class="wp-block-paragraph">A final direction attracting a significant amount of attention in the “what models learn” category was that of adversarial examples, which are data points which have semantically meaningful features corresponding to one category, but less semantically meaningful features which bias a model’s prediction in a different direction &#8211; for example, a photo that looks like a panda bear to humans but which contains noise that makes a model predict it to be a tree. Recent work in adversarial training has made progress in making models more resilient to such adversarial examples, and there were a number of papers at NeurIPS in this vein. I also had a very interesting conversation with Dimitris Tsipras, who was a coauthor on <a href="http://papers.nips.cc/paper/8307-adversarial-examples-are-not-bugs-they-are-features.pdf">this paper</a>, which found results which suggest that image classifiers may use some less-robust features for classification, which can be perturbed to generate adversarial examples without modifying the more robust features which humans primarily focus on. This is an emerging area of investigation and the literature is worth a closer look.</p>



<p class="wp-block-paragraph">All in all, it appears that the community is spending considerable effort in making models more robust and generalizable for use in the real world, and I’m excited to see what further fruit this bears.</p>



<h2 class="wp-block-heading" id="efficiency">Efficiency</h2>



<p class="wp-block-paragraph">As the power and applicability of deep learning grows, we are seeing a transition of the field from the 0-to-1 phase, in which the most important results have to do with what is or is not possible at all, to a 1-to-n phase, in which tuning and optimizing the techniques previously found to be useful becomes more important. And just as the deep learning revolution had its underlying roots in the greater availability of compute and data, so too are the most prominent directions in this area which I saw at NeurIPS concerned with improving the data-efficiency and the computational efficiency of models.</p>



<h5 class="wp-block-heading" id="data-efficiency">Data efficiency</h5>



<p class="wp-block-paragraph">Ultimately, deep learning depends on large amounts of data to be useful, but collecting this data and labeling it (for supervised approaches) are typically the most expensive and difficult stages of applying deep learning to a problem. A number of papers at NeurIPS had to do with reducing the severity of this issue. Many had to do with self-supervised learning, in which a model is trained to represent the underlying structure of a dataset by using implicit rather than explicit labels, e.g. predicting pixels of an image from neighboring pixels or predicting words in text from adjacent words. Another approach which a number of papers dealt with is semi-supervised learning, where models are trained on a combination of labeled and unlabeled data. And finally, weakly supervised learning has to do with learning models from imperfect labels, which are cheaper and easier to collect than perfect or almost perfect ones. Chris Ré’s group at Stanford, with their Snorkel project, are prominent in this area, and had at least one paper on weakly supervised learning at NeurIPS this year. This also falls under the “systems for ML” category, mentioned in the next section.</p>



<p class="wp-block-paragraph">Another prominent direction having to do with data efficiency (and also connected to representation learning) is that of meta/transfer/multi-task learning. Each of these approaches seeks to have models efficiently learn representations which are useful across tasks, thereby increasing the speed and data-efficiency with which new tasks can be tackled, up to and including one- or even zero-shot learning (learning a new task from a single example, or no examples at all). One interesting paper among many on these topics was <a href="https://arxiv.org/pdf/1912.03820.pdf">this one</a>, which introduces an approach to trading off regularization on cross-task vs. task-specific learning in the meta-learning setting.</p>



<p class="wp-block-paragraph">Another direction in data efficiency which I noticed prominently at NeurIPS had to do with shaping the space within which models learn to better reflect the structure of the world within which they operate. This can broadly be thought of as “stronger priors” (although it seems the term “priors” itself is being used less frequently). Essentially, by constraining learning with some prior knowledge of how the world works, data can be used for learning more efficiently within this smaller space of possibilities. In this vein, I saw a couple of papers (<a href="http://papers.nips.cc/paper/8396-scene-representation-networks-continuous-3d-structure-aware-neural-scene-representations.pdf">here</a> and <a href="http://papers.nips.cc/paper/9331-geometry-aware-neural-rendering.pdf">here</a>) improving models’ abilities to learn representations of the 3D world through approaches informed by the geometric structure of the world. I also saw a couple of papers (<a href="https://arxiv.org/pdf/1906.07343.pdf">here</a> and <a href="http://papers.nips.cc/paper/8825-learning-by-abstraction-the-neural-state-machine.pdf">here</a>, both from folks at Stanford) which use natural language to ground their representations of what they learn. This is an intriguing approach because we ourselves use natural language to ground and communicate our perception of the world, and forcing models to learn representations mediated by our languages in a sense imposes real-world priors upon the models. A final paper I’d mention in the category of priors as well is <a href="http://papers.nips.cc/paper/8777-weight-agnostic-neural-networks.pdf">this one</a>, which showed surprisingly good performance on MNIST of networks “trained” by architecture search alone &#8211; while this may not be immediately applicable, it is suggestive of the degree to which picking network architecture carefully (i.e. in a way that reflects the structure of a task) can make the learning process faster and cheaper.</p>



<p class="wp-block-paragraph">One final direction relevant to data efficiency is that of privacy-aware learning. In some cases (and likely more to come in the future), data availability is bottlenecked by privacy constraints. A number of papers I saw, including many in the area of federated learning, dealt with how to learn from large amounts of data without compromising the privacy of the people or organizations from which the data originated.</p>



<h5 class="wp-block-heading" id="compute-efficiency">Computational efficiency</h5>



<p class="wp-block-paragraph">As well as data efficiency, efficiency with regards to computational resources &#8211; i.e. compute and memory/storage &#8211; was also a prominent direction of many papers at NeurIPS. I saw a number of papers having to do with the compression of models and embeddings (the representation of the data used by models in certain settings). Shrinking models and embeddings/representations of data reduces both computational and storage requirements, allowing more “bang for the buck”. I also saw some interesting work in biologically-inspired neural networks, like <a href="http://papers.nips.cc/paper/8465-neural-networks-grown-and-self-organized-by-noise.pdf">this paper</a> from Guru Raghavan at Caltech. One motivation in this area is that while there will be certain limits to how many matrix multiplications and additions can be performed per dollar/second on general-purpose hardware to push the capabilities of modern deep learning, it may be possible to use special-purpose hardware which more closely approximates the functions of biological neurons to achieve higher performance for certain tasks. I heard a combination of curiosity and skepticism around biologically-inspired approaches from fellow NeurIPS attendees: this is an area to watch for the 10+ year horizon.</p>



<h2 class="wp-block-heading" id="directions">Directions and applications&nbsp;</h2>



<p class="wp-block-paragraph">Finally, while at NeurIPS I also found it very interesting to get a feel for the higher-level trends in various subfields of AI/ML and a feel for the different applications now possible, or becoming possible, thanks to recent advances in research. This section is more of a smorgasbord than a narrative; skip around as interest dictates.</p>



<h5 class="wp-block-heading" id="graph-nns">Graph neural networks</h5>



<p class="wp-block-paragraph">One area which I should mention seeing a number of papers around is that of graph neural networks. These networks are able to more effectively represent data in settings with graph-like structure, but as I know very little about this direction personally, I’ll instead refer interested readers to the <a href="https://nips.cc/Conferences/2019/ScheduleMultitrack?event=13172">page</a> of the NeurIPS workshop on graph representation learning as a starting point into the literature.</p>



<h5 class="wp-block-heading" id="rl">Reinforcement learning and contextual bandits</h5>



<p class="wp-block-paragraph">Another area in which I saw an absolutely tremendous amount of work was that of contextual bandits and reinforcement learning (RL). A few approaches which I saw a number of papers in were hierarchical RL (related to representation learning) and imitation learning (in a sense, setting priors for models through human demonstration). I also saw a number of papers dealing with long-horizon RL, in line with recent success in RL tasks requiring planning further into the future, e.g. the game Montezuma’s Revenge. A number of papers also had to do with transferring from simulation to the real world (sim2real), including OpenAI’s striking demonstration of teaching a robotic hand to solve a Rubik’s cube in the real world after training in-simulation. I also talked to Marvin Zhang from Berkeley about <a href="https://drive.google.com/file/d/1YJ0cg-WnJMWwzhZkh9y3J3QA5hduqM5X/view">a paper</a> he coauthored in which a robot was trained on videos of human demonstrations &#8211; “demonstration to real” rather than “simulation to real” learning.&nbsp;</p>



<p class="wp-block-paragraph">However, it is important to note that in practice, RL for the real world, i.e. hardware/robotics, is still not quite there. RL has found great success in settings where the state of the problem is fully representable in software, like Atari games or board games like Go. However, generalizing to the much messier real world has proved more difficult &#8211; even the OpenAI team behind the Rubik’s cube project spent 3 months solving the problem in-simulation and then almost 2 years getting it to generalize to a real robotic hand with a real Rubik’s cube &#8211; and even then, with far less than 100% reliability. It will be interesting to see how quickly new approaches to RL can square the circle of generalizing to the real world. I had a great conversation with Kirill Polzounov and Lee Redden from Blue River about this &#8211; they presented a <a href="https://drive.google.com/file/d/18dLTjFt5fCXoYRjRSytunh876Y2oVFm4/view">paper</a> on a plugin they developed for OpenAI gym allowing people to quickly test RL algorithms on real-world hardware. I’m excited to see how quickly “RL for the real world” progresses &#8211; if we see an inflection point like the one vision hit in 2012, the implications for robotics could be tremendous.</p>



<h5 class="wp-block-heading" id="nlp">Natural language processing</h5>



<p class="wp-block-paragraph">Another area worth mentioning is NLP (natural language processing), in which I’ve done some work personally. The Transformers/transferable language model revolution is still bearing fruit, with a number of papers showing good results leveraging those techniques. I was also intrigued by a <a href="https://papers.nips.cc/paper/9689-legendre-memory-units-continuous-time-representation-in-recurrent-neural-networks.pdf">paper</a> that claimed unprecedented long-horizon performance for memory-augmented RNNs. It will be interesting to see if the pendulum swings back from “attention is all you need” back to more traditional RNN approaches. It’s also worth noting that NLP is starting to hit its stride for real-world applications. I have a few friends and acquaintances working on startups in the field, including Brian Li of Compos.ai, whom I ran into at NeurIPS. I also enjoyed peeking into the workshop on document intelligence &#8211; it turns out NLP for the legal space is already a multi-billion dollar industry! Broadly speaking, natural language is the informational connective tissue of human society, and techniques to apply computational approaches to this buzzing web of information will only grow in the future.</p>



<h5 class="wp-block-heading" id="sysml">SysML</h5>



<p class="wp-block-paragraph">Another area I’ll treat briefly, from personal ignorance rather than unimportance, is that of SysML &#8211; i.e. systems for ML and ML for systems. This is an exploding field, as evidenced by the numerous papers presented at NeurIPS and the workshops in the field. One particularly interesting talk was the one Jeff Dean gave at the ML for systems workshop &#8211; definitely worth a watch if you can find a recording (please leave a comment if you do). He and his team at Google managed to train a network to lay out ASICs much more quickly than human engineers could, and met or even surpassed the performance of ASICs laid out by humans. A number of other papers also showed compelling results in optimizing everything from memory allocation to detecting defective GPUs with the help of deep learning. A number of papers also addressed the “systems for ML” direction, such as the Snorkel paper mentioned above.</p>



<h5 class="wp-block-heading" id="generative">Generative models</h5>



<p class="wp-block-paragraph">Generative models have reached a stage of significant maturity and are now being used as a tool for other directions as well as being a research direction in their own right. The performance of the models themselves is now incredible, with models like BigGAN having previously established a photorealistic state of the art for vision, and I saw a number of papers yielding unbelievably good results in conditional text-to-image generation, video-to-video mapping, audio generation, and more. I’ve been thinking about a number of downstream applications of these techniques, including some in the fashion industry and visual and musical creative tools, and I’m looking forward to seeing what emerges in industry in the years to come. Applications of generative models in other fields of machine learning has also been interesting, including fields like <a href="http://papers.nips.cc/paper/9127-deep-generative-video-compression.pdf">video compression</a> &#8211; I talked to some folks from Netflix about this, as it may prove useful for reducing the bandwidth load on the Internet from video. (Netflix and Youtube alone use something like ⅔ of the bandwidth in the U.S.) Generative models are also being used in sim2real work in robotics, as previously mentioned.</p>



<h5 class="wp-block-heading" id="misc">Miscellanea&nbsp;</h5>



<p class="wp-block-paragraph">Finally, for the sake of completeness I’ll mention a few more areas which I witnessed smaller bits of. Autonomous driving is still seeing a large and heterogenous amount of work. It seems that we’re settling into a state of incremental improvement, where both research and deployment of self-driving is going to happen in fits and starts over the next several decades (e.g. local food delivery with slow, small vehicles and truck platooning are easier problems than autonomous taxis in cities, and will likely see more commercial progress sooner). On the other hand, deep learning for medical imaging appears to be maturing as a field, with numerous refinements and applications still emerging. Finally, I was also intrigued by a paper in deep learning for <a href="https://arxiv.org/pdf/1906.01629.pdf">mixed integer programming</a> (MIP). Traditional “operations research” style optimization like that which can be framed as MIP problems drives tremendous economic value in industry, and it will be interesting to see if deep learning proves to be useful alongside older techniques there as well.</p>



<h5 class="wp-block-heading" id="conclusion">Conclusion</h5>



<p class="wp-block-paragraph">Modern AI/ML, largely powered by deep learning, has exploded into a large and heterogeneous field. While there is some degree of unsubstantiated hype about its possibilities, there is also plenty of genuine value to be derived from the progress of the last 7+ years, and many promising directions to be explored as the field matures. I look forward to seeing what the next decade brings, both in research and in industrial applications.</p>



<p class="wp-block-paragraph"></p>



<p class="wp-block-paragraph"><em>Thanks to Shengjia Zhao and Isaac Sheets for helping edit this essay.</em></p>]]></content:encoded>
    </item>
  </channel>
</rss>
