Very Impressed - 270/270 in my tool usage

#22
by Jakaryun - opened

I'm running it on an RTX PRO 6000(96GB) NVFP4 on llama.cpp operating on my custom harness with 270 tools via 3 external MCP and 1 internal for the harness(~174 tools). One MCP is a ticket-based shop + ERP system and works amazingly on generating invoices and tickets, scheduling, with contact lookup, price checking and many more with 93 tools. Another MCP server is for phone system management, it does great daily/weekly call reports. then another mcp for managing docker containers. This is the first time I have ran a local model and achieved a perfect tool call benchmark. Qwopus 27B Q8 came close at 190-199 out of 209 at the time(cant recall benchmarked a while ago). I will say Laguna does seem to pause and verify things more than I expected when doing tasks but it's actually been great responses. I do wish it had vision but worth not having it to run 512k context window. In bigger tasks, I noticed it would finish about 1/5th the time that Qwopus 27B Q8 took. Updated didn't realize it got a 100% on 270 tools, was 209 before I added more internal tools last time I ran it on Qwopus.

Jakaryun changed discussion title from Very Impressed - 209/209 in my tool usage to Very Impressed - 270/270 in my tool usage

Sign up or log in to comment