Skip to main content
Aggregate Semiconductor Engineering 芯片半导体 27 Aug 2026 - 16:01

Agentic AI Success Relies On Excellent Human Scaffolding

RSS 官方收录 · 可信分层展示

关键摘要

Key Takeaways: Long-running autonomous multi-agent systems have progressed significantly in the last year, but gaps remain in the full design flow.…

  • Companies and customers are starting to see positive ROI, with agents …
  • To conserve resources, human engineers need to define the ontology (st…
  • Experts At The Table: Semiconductor Engineering sat down to discuss re…

摘要引擎:抽取

正文提要

Key Takeaways:

  • Long-running autonomous multi-agent systems have progressed significantly in the last year, but gaps remain in the full design flow.
  • Companies and customers are starting to see positive ROI, with agents providing a step function in design capability rather than incremental improvement.
  • To conserve resources, human engineers need to define the ontology (structured knowledge, relationships, and concepts) of a specific domain, as well as the agentic harness, which is software infrastructure around the model, consisting of tools, memory modules, execution loops, and guardrails.

Experts At The Table: Semiconductor Engineering sat down to discuss recent developments and challenges when using agentic AI for chip design, with Matt Graham, senior group director of verification software product management at Cadence; Harrison Balistreri, head of business development and strategic partnerships at ChipAgents; Alexander Petr, senior director and portfolio manager at Keysight EDA; Sathish Balasubramanian, head of product for EDA AI at Siemens EDA; and Anand Thiruvengadam, executive director and head of AI product management at Synopsys. This roundtable discussion was held behind closed doors at the recent Design Automation Conference. What follows are excerpts of that discussion.

Fig. 1: L-R: Synopsys’ Thiruvengadam; Keysight’s Petr; Cadence’s Graham; ChipAgents’ Balistreri; and Siemens’ Balasubramanian. 

SE: A lot has changed in the past year in terms of agentic AI capabilities. A year ago, people would say, ‘Agentic AI is going to be useful for this, but it can’t do this.’ Now, what people said it probably couldn’t do, such as full-flow design, it’s now doing to some extent. The points where humans are definitely needed are getting smaller and smaller. What are your thoughts on where the design industry is at with this?

Balasubramanian: Compared to last year, we are talking more about orchestration of different flows or sub-agents. Customers are getting a handle on what it is. I’m seeing a lot more automation on the front end — design, mainly RTL generation and sign-off, or debug. They are all creating sub-flow agents. The end goal is an end-to-end, push-button flow. Give us a spec, get a GDS II or a chip out. That’s the main end goal. I don’t think we are there yet, but customers are using agentic AI in production for specific pain points with a lot of manual intervention and small manual tasks. That’s getting much more automated with agents.

Balistreri: We’ve progressed quite a bit, but you have to deliver the agentic AI in a certain way to get that progression. We think a lot about full-flow autonomous outcomes, and we’ve delivered this for specific engineering flows. Think of us as push-button coverage closure with multi-agent swarms. This is quite different than the point acceleration that you saw last year, with an agent for this, an agent for that, and generic use of cloud flow. To get true full-flow autonomous outcomes, you need several layers. You need organizational context. You need to understand the flow specific to that customer and how they do that engineering flow, and map how you use agents to that flow. You need an agentic harness tuned for these chip-designed tasks; then you can co-optimize a model with that harness to run these flows efficiently and specifically to the harness. We’ll come back to this idea of critical errors, too. You cannot have a full flow, fully autonomous [system] without certainty that you’re not introducing critical errors. What we’ve progressed to now is delivering this full-flow autonomous outcome without critical errors. But it takes a lot, and we’re working very closely with each customer to map it.

Graham: One thing that’s the same from last year to this year at DAC that I’m really excited about is that we do see a very vibrant startup community in EDA. A few years ago, it was slim pickings, and this is fantastic. This is great for the industry in general. This is great for innovation, and historically this is how EDA has grown — both organically with the big guys and through startups, regardless of whether their technology advances the industry or gets absorbed by the big guys through acquisition. I’m happy to see that continue. As far as what’s different, last year we talked about agentic AI almost exclusively in the front end — RTL, test bench, and things like that. Certainly, that has advanced, but we’re now seeing it in analog and mixed-signal. We’re seeing it in full custom. We’re seeing it through digital implementation, even into circuit board and packaging. Last year’s push-button flows were primarily in the front end. This year we’re talking much more about genuine end-to-end flows and even expanding into multi-physics and beyond just the IC itself, or beyond just the ones and zeros. We’re seeing the entire industry grow, and agentic AI is now moving into the full chip-and-system flow.

Petr: The biggest change was around February this year, when OpenAI released a reasoning framework. Since then, we have been talking about super agents or agent swarms. Last year, we talked about building a single agent that can understand a unique domain. That is very time-consuming because it means tailoring engineering knowledge to a framework that understands what needs to be done. From what I see, a lot of effort is still in the digital domain, where you still have code as a base. The agents are horrendously good at coding. Period. They understand code. They speak Python naturally. As soon as you go into a domain where people still use UI and a mouse to do most of their work, this is where we still struggle. To Matt’s point, right now what we’re enabling is multiphysics. We’re trying to enable all our core simulation technologies, all the glue between the different tasks, which are still very engineering-heavy, to get to the multiphysics package solution. Everything we see nowadays — if you look where the investment goes, the CapEx goes —is multi-chip technologies. It’s the AI stack. It’s photonics. It’s not just a single chip anymore. It doesn’t scale. Moore’s Law on that side is flat, so the scale comes from integration. Enabling all those pieces that enable integration is the hard part now. We see a lot of design agents that go beyond RTL and other things. We see engines that were ultra slow in the past being enabled through AI technologies, and now we can put an agent on top and integrate it into an agentic framework. The call is still out on whether one will rule all, or if there will be multiple. A lot of people here are fighting over who will rule the world through super agents or frameworks. So are the startups. If you ask me what’s going to change next year, there will be an awakening at some point about how this will pan out. Will there be one, or will there be multiple, or will each company pick its own?

Thiruvengadam: From a broad industry point of view, there are multiple things happening. There’s marked acceleration in some of these trends that started a few years ago. Microsoft, Nvidia, and many other ecosystem players have essentially enabled four pieces of technology that have fueled some of the innovations that the others at this table are talking about. The EDA vendors have leveraged that. This year is pivotal because task-level automation is great right now. Workflow automation is great. But we are truly at the point where we see long-running agentic workflows or AI systems as a reality because we’ve demonstrated it. This is good for the broader industry because you can now think about fully autonomous execution with minimal human supervision — within human guardrails, as needed. We have end-to-end workflows that can run end-to-end in a secure environment. That’s happening as a reality. It’s a significant evolution. The second thing is business outcomes. We are talking about the technology, and it is great. Yes, long-running agentic workflows are here, but on the business outcome side we are starting to see that, as well. Last year, we talked about some very initial gains, early gains, early trials, in some very specific domains such as verification. Now we have fairly broad-based outcomes established by customers, including verification implementation. Analog, we’ve seen, as well. On the system side, with structures, CMT (common-mode transient), and thermal, we are seeing real, tangible gains for our customers. With these long-running capabilities, it’s not just point-level task automation. If you run a fairly complex end-to-end flow through a long-running system, you see the outcomes we expect from long-running agents, such as significant productivity gains and improved software quality. So technology has progressed, yes, but so has business — actual value outcomes for customers. We’re seeing positive ROI.

SE: What does the engineer need to do before deploying an agent? How is that changing the setup before a chip is designed, using the engineer’s preferred EDA tools?

Balasubramanian: Ontology is the key thing that an engineer needs to do right now. That is where we are seeing people spend more time. With agents, no one worries about scripting or automation. They’re thinking more about design and architecture. ‘How else can I improve a certain sort of architecture?’ Engineers are thinking more about, ‘What are my bounds on PPA?’ or whatever their metrics are, and what domain they are in. Then they’re figuring out, ‘What are all the different parameters that I want the agents to work on?’ Given that we talked about long-running agents and swarm agents, they’re letting the agents and LLM-driven automation explore a design space to find the best result and ROI for their particular design. The important starting point for the engineer is clear domain knowledge — not being constrained by the previous design where I want to add margin and so forth — but thinking about what really makes designs successful and how they fit into the bigger picture. After that, it’s all a matter of how you drive the agentic AI. We are seeing some companies use it in ways that deliver huge benefits. Instead of iterative design improvements, agents are delivering a big step function in implementation.

Thiruvengadam: In terms of the role of humans, there are two key things — humans will be specifying the design intent. That’s where specs come in. That’s a huge role that I still see happening. More important, with all this capable automation and long-running agents, is that humans are spending a lot of time on verification, as in verifiability and validating the outcomes. Those are the two big functions that will not go away. What I see is a marked shift toward human engineers spending a lot of time on verifiability, and that ties into how these systems are architected and designed. Verifiability can be done in different ways. A very intuitive user experience for verifiability is key. You need to present something that’s easy to digest and verify. Without that actual deployment, that’s an orthogonal point. Agentic flows have to be designed with the right checkpoints, user interface, and intuitiveness. All of that needs to be hand in hand. It’s not just somebody looking at a verbose report and saying, ‘Okay, what do I do next?’ Agents need to present outcomes in a way humans can quickly validate and move on. Otherwise, you’ll create yet another bottleneck.

Balistreri: I would frame it differently, because if you’re just producing things for humans to verify, you’re not delivering on the autonomous piece. That’s what is missing in ROI today for most agents and most flows. To deliver on the autonomous piece before engineers deploy a solution like ours, they have to trust the process’ outcomes. The gates aren’t put in the flows themselves. If you’re using a generic LLM, or you’re just piping an existing LLM into your solution, and that powers your agents, you’re using a very thin agent awareness in the agent loop that Anthropic or OpenAI has picked for you. We’ve put the trust in the agent loops themselves. We’ve reconstructed agent loops from scratch in a way that prevents critical errors with 100% certainty. When you deploy the agents, you know they will not produce certain classes of error. I want to share one example — waivers in your coverage model. We’ve found that most agents can decide to waive coverage when they shouldn’t, creating a massive human review burden. We’ve heard from one top-five customer at TSMC that their biggest pain point is when their engineers use Claude code, especially their junior engineers. They’re creating the equivalent of AI slop with all of these errors in it. But I’d say it goes even further into the harness you’re using and the agent loops themselves. In our full-flow autonomous coverage closure, we have to use gates built into the agent loops themselves to prevent you from waiving coverage on a cover bin, and to prevent you from saying you’ll never converge on coverage for this solution, or things like that.

Thiruvengadam: I completely agree. If you’re talking about fully autonomous long-running systems, then yes, the self-validation is built into the harness orchestration loops themselves. No disagreement there. The point is we are still early in broad-scale deployment of long-running systems — that’s the key. So verifiability will still be there. I’m not saying a fully autonomous end-to-end workflow, which you think of as a black box, will still have elaborate, laborious human checkpoints. That’s not the point. The point is that even if it’s a fully autonomous verification flow, which we have demonstrated, or a debug, or an implementation flow, critical checkpoints are still required for human engineers, and that’s the key because the trust factor isn’t there yet. At some point, it will hopefully evolve to remove all the guardrails and maybe have fairly discrete, less-frequent checkpoints.

Balasubramanian: That’s where the ontology matters. In terms of waivers, customers have been doing waivers for a long time. It’s nothing new for them, but we are trying to help them do it faster. That’s the whole idea for people who are doing tape-outs. For ontology, how are you going to tie up the harness with the existing ontology across the entire flow, not just one particular thing? For example, you can have coverage on one verification flow, but a separate IP team will be doing their own thing. A lot of new data and modifications are coming in, so you need a full picture. Ontology is very important. You can’t just be on one single domain and say, ‘Okay, I’m going to be just covering this one.’ It needs to touch everything, wherever all the key domains are that are going to be coming into that part, and also going downstream. That’s going to be key. On the verification side, I agree about AI slop. We call that design slop. The biggest challenge we see, even on the verification side, is it’s not just humans verifying. You still need the engines at some point. What we are seeing is customers get 10 scenarios coming in, and based on their domain knowledge, the verification team picks 5. But those 5 still need to go through the engine layer for sign-off because LLMs aren’t deterministic. You can add guardrails, skills, and everything else, but as you go down the stack toward the silicon, you really need to figure it out. That’s where the humans come into play to narrow down your scenarios, and then you’ve got to go through the sign-off engines and make it mathematically accurate, so that the chip doesn’t fail. At the end of the day, you’ve got to make the chip.

Graham: We can’t emphasize enough that the fundamental engines are still very deterministic and mathematically accurate. At some point, we need them because the cost of failure is too high. Whether it’s in verification, implementation, analog simulation, or signal integrity for the board and package, the integration between the AI and where [the engines] meet is still critical. That integration plays a major role in setup and preparation. Taking a slightly different direction, another difference I’m seeing more than last year — in terms of preparation before you deploy AI — is that token budget has suddenly become a meaningful line item for all of our customers and, I’m assuming, for all other companies as well.

The post Agentic AI Success Relies On Excellent Human Scaffolding appeared first on Semiconductor Engineering.

打开官方原文 站点原文页 可信分区 本信源更多 今日简报 分享图 RSS 稍后再看列表