从语音输入到端到端税表
00:00:00–00:11:16Insight
- 口述比书写快,也减少人在写作时无意删去模型可能需要的背景。
- 自主完成不等于无人查看:交付要列出重大决定、假设和复核点。
- 好的输出会优化审阅成本,类似拆分清晰、附解释的代码提交,而不是一份难以检查的巨大结果。
Humans are already used to working with nondeterministic systems. It's just those systems are normally their co-workers, not their computers. And in many ways, like companies and processes is all about how do you design a system for nondeterministic entities to coordinate together to solve a problem. And once you realize that, it's like, well, now it's like kind of agent design. Let's say you have 100 evals. Great. They all pass. It looks good. Are you confident that that now
generalizes to the real world to production?
And our answer has been no. Even if you got it right a 100 out of 100 times, if a person is just getting it right because they're going to Wikipedia, the accounting firm wouldn't hire them and so they shouldn't hire us either. You'll see people like freaking out over a code file that isn't abstracted properly and yet their context is like total. The English is more precious because the English affects the performance. The code does not affect the performance. Hi, I'm at from
Mark. Welcome to the Mad Podcast. Everyone is b uilding AI agents, but outside of coding, most still can't do real work reliably. My guest today is an AI builder at the forefront of cracking that problem. Mitch Troyanovski is a co-founder of Basis, a unicorn AI company whose agents run autonomously for hours, sometimes days, and are already able to complete very complex tasks like preparing entire tax returns end to end. This is a true reference episode on how to build long h
orizon autonomous agents where Mitch shares tons of lessons he learned along the way. Please enjoy my conversation with the deeply insightful Mitch Troyki.
I want to start with a scene. Uh, as I was uh prepping for this, I uh came across a video by our friend Stephanie Palado at the information and she was describing um uh the experience of walking into the basis office and seeing a bunch of people uh whispering very quietly into microphones. So you know maybe for the top AI builders or people who live on X you know second by second this may be already something that everybody understands but I think for the vast vast majority o
f people like just describe what you guys are doing whispering into those microphones. Yeah, I I think maybe the best piece of advice for not for building agents but for working with AI in general is that you need to give it as much context as possible because it by definition is always missing context in some way. Speaking is just so much faster than writing things down. And in fact, when you try to write things d own,you are actually essentially uh trying to summarize all t
he crazy thoughts in your head. And so that's why it takes a lot of time. And it's useful for you or I because it's rude if somebody just blabbered and sent that to you as a as a Slack DM, but to an agent, they don't care. Um they uh it's actually they would prefer it. So it it becomes much more productive to be able to like whisper your thoughts, you know, because you don't want to be shouting. Um uh and you have these microphones now that
主持人 allow you to whisper very very quietly uh and still pick up with full fidelity. So that's why that's why we have it. And sometimes people see it and they think it's a little weird when they join the company, but after after a couple after, you know, a month or so, they're it's like they they can't go back. So,
嘉宾 and uh so you whisper into what? Into cursor or into
主持人 Yeah. into whatever people use. I mean, we different people use different things, but yeah, codeex or cloud or cursor or whatever people use and not just engineering, right? Like all the functions if you're trying to get something done, if you're trying to describe what you want and all these things.
嘉宾 Okay, great. All right. So uh what I'm hoping to uh do today is uh is a bit of a reference conversation on uh all things around building long horizon agents that's in part based on um a great thread that you had on X and perhaps more importantly a new open source uh project that you just released in collaboration with Brain Trust. We're going to talk about all all about this but maybe for contextual awareness basis in two or three sentences how would you describe it?
Yeah, basis builds agents to do accounting work end to end and accounting is difficult. Um it's not something that is just purely text in text out and so it requires the ability for like AIs to be able to you know perform lots of actions over long periods of time and actually um you know be coherent over that period of time to get to outcomes that are that are good. Um and that's why we've always you know been very focused on how do you really build agents that can scale to do that work. And did you pick accounting uh because uh of how interesting that was from a uh agent building uh perspective or the other way around?
主持人 Good question. Uh it's probably uh the other way around, but I I do think it is actually quite interesting from an agent perspective. Accounting is interesting for a lot of reasons. It is uh one one of if not the largest knowledge work profession in the country. Um there are over three million you know combined kind of accountants uh in the country. And what I think is so cool about accounting actually is that most people don't really think about accounting. They don't think like, oh yeah, why does that even there?
You know, probably most listeners have never thought like why does it even exist?
And um I know we're going to talk about agents maybe quickly 30 seconds just to convince everyone how cool accounting is. Uh if you think about the real world, uh so much stuff happens, academic activity, right?
Like you know, I was just drinking a water bottle there. like you know someone um uh that bottler had to choose to like go uh uh buy from that factory or that supplier um or decide to open you know some additional uh store hire a salesperson and these are all economic decisions that stem from understanding the real world what's in the real world you know money moves hands someone signs a contract someone delivers the inventory it's like all these events that occur and so much of modern capitalism relies on the ability of all these actors to make decisions on these events right?
Like the CEO of that company, the IRS obviously to decide how much to tax, the uh bank to lend credit, investors, right?
All these people, they care about the real world, but they can't understand it because it's gigantic and it involves all of this unstructured and, you know, difficult to parse information. And accounting is actually the art of compressing all of that into something that is structured that now people can look at and understand and make decisions. So something about accounting, you could argue in a meadow way, is kind of like an intelligence over the economy. um uh because it i
s really a compression activity of all the information that exists. So I think it's a very cool problem to kind of think about.
嘉宾 This is still the same thing for contextual awareness. So we're going to talk about long um horizon agents. What is a I guess what what is a long horizon part these days? So that that keeps evolving. Uh what falls in that category?
主持人 Yeah. So maybe I can give my quick definition of an agent just I know maybe probably everyone knows at this point but uh feel like that's a gotcha question I like to ask in interviews. Um uh I tend to think of uh an agent as an AI or like some inference that occurs that has the agency to go and make uh decisions to to do different things. And so by definition it's a spectrum because you can have varying degrees of agency, right?
Like you're constrained by whatever environment you're placed in. And I think long horizon again is a spectrum where you are granting the agent the agency to make decisions that allow it to be coherent for longer periods of time, right?
So let's say that you were asking an agent to go and you know look up uh the weather for you. It might be an agent in the sense that it has the agency to decide what tool to call or what Google search to put in. But it doesn't need to do much work to be coherent over a period of time because it's uh um you're just getting the weather. But if you're asking an agent to say um go perform you know an entire uh feature you know um like implement some feature in your in your repo or asking it to go and you know make a big Excel workbook. Now suddenly it might have to operate for you know longer than a minute. We're talking 10 minutes, 20 minutes, 30 minutes, and potentially even much longer than that. And once you're starting to get into those scales, you start running into the fundamental limits of how LLMs work in which I always like to say LM have very large working memories and by default no short-term or long-term memory. And so you have to leverage these strengths of the LM to make up for the fact that they don't have long good uh or actually any real short-term or long-term memory um by and we can talk more about it by using harnesses and you know all these kind of uh advancements to allow them to be coherent over a period of time. So I think once you start getting into the art of trying to get it to be coherent because you're going past the the you know the amount of working in memory it has I'd probably say that's when you're starting to get into what I'd call long horizon.
嘉宾 Great. And still to the conversation, what we're talking about here is uh autonomous
主持人 uh agents. I'm curious maybe uh just as an example what autonomous means in the context of basis. So I read that you guys can now have agents that handle um end to end tax returns.
嘉宾 So may maybe walk us uh at a high level through what that looks like in terms of steps along that takes. uh what does an agent do conceptually
主持人 when you are really autonomous or kind of you know doing something over really long horizon say doing a tax return and hand that means that you uh have a lot of information that is needed to do the work and you have the tools to go and get potentially more information. So, let's say imagine you're doing a complicated 1065 and you have all of the different um uh you know K1s, W2s, other documents, 1099s, whatever you need um from the uh from the company. Uh uh and then you als
o potentially depending on what you're doing, you might have the trial balances already. So, that tends to be what you need to actually start a tax return or you're working with like books that aren't even done yet. Um, and the agent then has to actually go and figure out based on all this different stuff, how is it going to tackle it and what's it going to be able to do?
And that's where you start getting into um, you know, some stuff about like what the behavior should be that we can talk about about well, what does good practice look like to say get to a solid solid uh, set of trial balances. What does good look like in order to properly extract out, you know, the K1s and the K3s so that you can be confident in their outcome. Um, and for it to be autonomous. It means it's not going back to the user and saying, "Hey, like is this right?
Is this right?
I need this. I need this." It's like starting the job to I'm done. Um, and I'm done does not mean I'm done. You click a button like no one looks at it. It's actually the opposite of that. It's much closer to what you can imagine a preparer doing or or maybe like a first pass or junior engineer or something of I'm done. Here were the big decisions I made. Here were my assumptions. Here were the different things you need to look at. Let's go and review together. Right?
And if you think about somebody say in engineering, you know, an engineer handing you a PR, nobody likes being handed a thousand line PR, they're like, it's done. I promise. It's like you don't want to review that. But if you instead handed somebody, you know, a great stack that was like that was like properly split out and you could understand very easily, hey, here is exactly what this change is and this diff and I made this big architectural assumption here and here's why I made that change. and you can like optimize not just for getting the work done but for making it easy for your reviewer to understand the decisions that you made. Uh and that's obviously very true in software engineering and it's actually true in I think most professions uh and especially accounting which we can kind of get more into. And so to me that's I think what it means to be to be autonomous. So I thought what would be fun and helpful uh for people listening to this would be to spend a few minutes on I guess the history of agents like we've all heard over the last two to three years so many different things so many different terms some projects that work some projects that didn't work so I think it would be helpful to just like go back in time just a little bit what in AI may feel like a prehistory but in reality is like what three years ago four years ago so maybe starting uh in 2022 with the react uh framework. So not the software engineering but like reasoning and and acting which I believe was a paper in 2022 that fundamentally said this agents are combination of like reasoning and acting which you you just alluded to. The fundamental question is that is that still largely what's what's happening?