AI Agent for Business: We Tested Hermes Agent — The Result Surprised Us

An AI agent for business can do more than generate text: it can plan steps, use tools, and work with files. Our Hermes Agent test found that the greatest value does not come from unlimited autonomy, but from restricted workspaces, suitable models, and reviewable tasks. The result is capable, yet far from maintenance-free.

Why did we test Hermes Agent in the first place?

The conversation about enterprise AI has moved beyond writing emails, summarizing meetings, and generating marketing copy. Companies are now asking a more demanding question: Can an AI system prepare and execute actual work instead of merely discussing it?

That is where AI agents enter the picture. An agent can be authorized to browse websites, inspect files, run terminal commands, use internal knowledge, call APIs, and pursue a goal across several connected steps. The change may appear technical, but it creates an entirely different operating model. Once an AI system can take action, an incorrect assumption is no longer limited to a poor paragraph. It can become an operational incident.

This matters to German midmarket businesses because AI adoption is rising rapidly. Bitkom e.V. (https://www.bitkom.org/) reports that 41 percent of German companies with at least 20 employees already use AI. Among companies that use it, 77 percent say AI has improved their competitive position.

Those figures explain why many executives are no longer satisfied with a general-purpose chatbot. They want systems that can collect information, process documents, connect knowledge from separate sources, and reduce the manual effort behind recurring business tasks.

Hermes Agent attracted our attention because it attempts to provide a broad, open-source environment for exactly this type of work. The project is developed by Nous Research (https://nousresearch.com/) and combines language models with tools, workspaces, memory, reusable skills, browser capabilities, terminal access, and multiple execution options.

On paper, that resembles a general digital employee. In practice, our test produced a more useful and more nuanced result.

AI Readiness Assessment by KrambergAI

Assess where AI can create real value

The KrambergAI AI Readiness Assessment helps companies identify suitable AI use cases, evaluate process readiness and define realistic next steps for structured implementation.

Structured assessment · Practical prioritization · Made in Germany

How did we structure our Hermes Agent test?

Our hands-on evaluation used Hermes Agent v0.18.0. The project is evolving quickly, so later releases may contain different menus, setup steps, tools, or security controls. The broader observations about model performance, permissions, workspaces, and operational ownership remain relevant.

We ran Hermes on several Apple Silicon systems, including a MacBook Air, a MacBook Pro, and a more powerful Mac Studio. This allowed us to compare smaller local models, larger local models, and cloud-based models without changing the overall agent environment.

Terminal operations were routed through Docker. Hermes did not receive unrestricted access to the entire user account. Instead, it worked inside selected directories created for the test. This decision turned out to be more important than many of the agent’s visible features.

An AI agent that can execute shell commands, modify files, install packages, and open browser sessions should not begin with the same access rights as the person operating the computer. A dedicated workspace limits what the agent can reach and makes it easier to inspect what changed after a task.

We focused on practical business work rather than demonstration scenarios. The test included market and technology research, local document analysis, structured report generation, file editing, small scripts, technical concept development, and attempts to convert recurring procedures into reusable skills.

We intentionally avoided connecting messaging platforms that could make the agent permanently reachable from outside the test environment. External channels may become useful later, but they are not essential for proving business value. They also introduce additional identity, authorization, and audit questions before the underlying use case has been validated.

What makes Hermes Agent different from a standard AI chat?

A standard AI chat usually works within the current conversation. It answers a question, drafts content, explains a concept, or summarizes information supplied by the user. Hermes Agent can go further by creating a plan, selecting tools, inspecting intermediate results, editing files, and initiating additional steps.

The differentiator is not a single language model. Hermes can work with different local and external models. Its practical value comes from the operating framework that connects the model with tools, workspaces, memory, skills, and execution environments.

CriterionStandard AI chatHermes AgentDeterministic workflow automation
Primary purposeGenerate answers and contentComplete multi-step tasksExecute predefined process steps
Handling variationResponds to new promptsCan adapt its approach during a taskRequires defined rules and exception paths
Tool accessLimited by the interfaceFiles, browser, terminal, skills, and other toolsPreselected systems and APIs
Reusable knowledgeConversation or project contextPersistent memory and procedural skillsRules, mappings, and process logic
AuditabilityOften limited to conversation historyDepends on logging and configurationUsually documented through workflow instances
Best suited forDrafting, explanation, and analysisVariable knowledge work across multiple stepsRepeatable transactions governed by stable rules
Typical riskIncorrect or incomplete outputIncorrect output followed by an unwanted actionAn incorrect rule is executed repeatedly

The comparison also explains why Hermes should not be treated as a replacement for ERP workflows, business process management, or conventional automation. When an order must always be created, checked, and posted according to fixed rules, deterministic software is usually the more dependable option.

Hermes becomes useful when the path cannot be fully predefined because information must first be found, interpreted, compared, and transformed into an appropriate deliverable.

What worked unexpectedly well during the test?

The surprising part was not that Hermes could write text or search the web. Established AI interfaces already perform those tasks. The agent became more interesting when one request required several different kinds of work.

Hermes could research sources, collect findings in a working file, identify missing aspects, and turn the material into a structured draft. During technical tasks, it could inspect files, execute commands, analyze error messages, and attempt a revised approach. That sequence felt substantially closer to everyday knowledge work than a single response in a chat window.

Separate profiles were also valuable. One profile could use a local model and a tightly restricted workspace, while another profile could use a stronger cloud model for more demanding reasoning. The result was not a forced choice between local AI and cloud AI. Different tasks could be routed according to data sensitivity, performance requirements, and available infrastructure.

The skill system also demonstrated practical potential. A recurring procedure can be stored as an instruction package so the user does not have to restate the expected workflow, reference files, and output format each time.

A skill should not, however, be confused with an approved business process. It remains guidance interpreted by a probabilistic model. A model may misunderstand a step, skip a verification, or apply an instruction in an unexpected context. Business use therefore requires versioning, testing, approval, and a way to roll changes back.

The strongest aspect of Hermes was ultimately not autonomy. It was the ability to combine models, tools, knowledge, and execution environments in a modular way.

Where did Hermes Agent reach its practical limits?

The first constraint was model capability. Smaller local models handled basic file operations, short analyses, and limited tool sequences reasonably well. As tasks became longer, they were more likely to forget an intermediate requirement, select the wrong tool, or misinterpret an error message.

Larger local models improved reasoning and tool use but ran noticeably slower on workstation hardware. Performance on the Mac Studio was acceptable for research and technical preparation, yet it remained behind strong cloud models. That difference matters when an employee is waiting for an answer during an operational process.

The second constraint was the surrounding technical environment. An agent can use only those tools whose dependencies are installed and functioning. Browser automation, integrations, local packages, and external services may require additional setup, credentials, and troubleshooting. A feature appearing in an interface does not guarantee that it is ready for dependable business use.

The third constraint was error propagation. A chatbot can provide a poor recommendation. An agent can take that recommendation, modify files, execute a command, or send data to another system. The same reasoning error therefore carries more operational weight.

This observation matches broader adoption patterns. McKinsey & Company (https://www.mckinsey.com/) reports that 62 percent of surveyed organizations are experimenting with AI agents. Moving from experimentation to repeatable use across business functions remains a much more demanding step.

Another limitation was operational visibility. During a long task, a user must be able to understand which tools were used, which sources affected the result, which files changed, and where the process deviated from the expected path. Without sufficient logs, the agent may appear productive while creating hidden review work.

What usually goes wrong when companies introduce AI agents?

The most common mistake begins with an oversized mandate. The agent is expected to answer email, prepare proposals, update CRM records, monitor competitors, coordinate appointments, and manage internal documents. That is not one use case. It is a collection of unrelated processes with different data, owners, exceptions, and risk levels.

The result is usually a prolonged configuration effort followed by inconsistent outcomes. Teams then debate whether the model, prompt, data, API, or workflow caused the problem. The agent becomes the technical location where unresolved process weaknesses accumulate.

A second mistake is starting with production credentials. A broad API token, an unrestricted user directory, or a service account with extensive permissions can make a demonstration easier. It also creates exposure that is unnecessary for an initial pilot.

A third mistake is treating technical functionality as business acceptance. An IT team can verify that Hermes reads a document, runs a command, and generates an output file. It cannot independently determine whether a proposal is commercially sound, a maintenance report is complete, or a customer response is professionally appropriate. A business owner must review the result.

Measurement is often weak as well. “The agent completed the task” is not enough. A useful evaluation should consider review effort, error types, missing sources, repeated deviations, recovery behavior, and whether another run produces a comparable result.

Gartner (https://www.gartner.com/) predicts that more than 40 percent of agentic AI projects will be canceled by the end of 2027 because of escalating costs, uncertain business value, or inadequate risk controls.

That forecast does not suggest that AI agents lack potential. It suggests that technology alone cannot compensate for a poorly selected process, undefined ownership, or an absent economic model.

How secure is a local AI agent for business?

Local operation is often treated as a synonym for security. That assumption is incomplete. A local model may keep prompts away from an external model provider, but the agent can still damage local files, expose information through logs, process malicious documents, or transmit data through an authorized tool.

The central security question is therefore not only where the model runs. It is what the agent is allowed to do.

Docker provided a useful starting point in our test. Commands ran inside a separate environment, and only selected workspaces were mounted. Reference material can also be exposed as read-only while generated files are written to a dedicated output directory.

That separation must be designed with care. A container provides limited protection when the entire home directory is mounted with write access. Credentials require the same discipline. Only secrets required for the specific use case should be available inside the agent environment.

A business deployment should also include approval gates for risky operations, separate service identities, audit logs, controlled updates, and recurring tests for prompt injection. Content retrieved from websites, emails, attached documents, or external knowledge bases must not automatically become trusted instruction.

Persistent memory and self-generated skills create another security dimension. Information stored after one session may alter later behavior. Companies need a reviewable record of what was learned, who approved it, when it changed, and how it can be removed.

Open-source access is valuable because the code can be inspected and adapted. It does not remove supply-chain risk, dependency management, configuration errors, or the need for security testing.

Which business use cases fit Hermes Agent?

Hermes is most useful when an employee currently moves between a browser, file system, spreadsheets, internal documentation, and several information sources. The agent does not remove every handoff, but it can take over a significant share of the preparatory work.

Research and market monitoring: Hermes can search approved sources, extract relevant material, organize findings by criteria, and prepare a report. A business owner then reviews the sources, interpretation, and conclusions.

Document preparation: Existing notes, templates, and reference documents can be used to draft project reports, statements of work, meeting packages, or technical documentation.

Internal knowledge work: Within a restricted document collection, Hermes can locate information, identify relationships, and prepare answers. Binding use requires maintained source documents, priorities between conflicting sources, and version history.

Technical assistance: The agent can inspect configuration files, draft scripts, run tests, and analyze errors. Production changes should continue through version control, peer review, and established deployment procedures.

Recurring back-office preparation: Hermes fits tasks that follow a recognizable pattern but involve different information each time. Examples include supplier comparisons, tender reviews, account preparation, market scans, and customer meeting briefs.

Content operations: An agent can collect source material, compare claims, assemble first drafts, and prepare localized versions. Editorial review remains necessary, especially when the content contains statistics, legal statements, product promises, or technical recommendations.

Hermes is less suitable for autonomous payments, binding contract decisions, employment actions, safety-critical controls, or unreviewed changes to production systems. In these areas, an agent should prepare, analyze, or recommend rather than make the final decision.

When is Hermes Agent the wrong choice?

Hermes is the wrong tool when a process is entirely rule-based, transaction-oriented, and already documented. An invoice with fixed validation steps, a standardized data transfer, or an order created under mandatory business rules does not need a language model that can improvise.

In such cases, agent behavior introduces unnecessary variation. Conventional workflow software can document each step, enforce exception handling, and repeat transactions predictably.

Hermes can also be excessive for very small tasks. When the only requirement is summarizing a document or drafting a short message, a standard AI interface is sufficient. Configuring workspaces, models, skills, tools, and approval rules would add more effort than value.

The system is also a poor fit when no one can assume technical ownership. Open-source software can reduce vendor dependency, but it does not eliminate operations. Updates, packages, model changes, access rights, logs, backups, and security controls remain the responsibility of the company running it.

A final warning concerns demonstrations. A successful one-time run proves that the task is possible under those conditions. It does not prove that the process will remain dependable after a model update, a website change, a modified document format, or a failed API request.

The decision should therefore be based on sustained performance and manageable operating effort, not on the most impressive individual session.

How should a midmarket company structure a Hermes Agent pilot?

A useful pilot begins with a business task rather than the software installation. The task should occur regularly, consume measurable effort, and produce an output that a subject-matter owner can review.

A weekly market briefing is one example. The agent receives selected sources, a reporting template, and a dedicated workspace. It performs research, organizes findings, and produces a draft. The company measures processing time, source quality, correction effort, technical failures, and repeated deviations.

Initial runs do not require production integrations. Exported files, synthetic records, or copied documents are sufficient to validate behavior. Direct write access to ERP, CRM, or document management should come only after the preparation task performs reliably.

The pilot also needs a documented baseline. This includes the selected model, system instructions, installed skills, authorized tools, mounted directories, credentials, and approval rules. Without that record, later behavior changes become difficult to diagnose.

Economic evaluation should cover the entire operating process. Model usage is only one cost component. Setup, maintenance, subject-matter review, error correction, testing, and updates also consume resources. Potential benefits include reduced research time, shorter turnaround, improved consistency, and better reuse of internal procedures.

A pilot succeeds when the agent reliably reduces effort within a bounded step. The number of actions performed by the agent is not a meaningful success metric on its own.

What is our final assessment of Hermes Agent?

Hermes Agent is not an instant digital employee that can be given a job description, credentials, and access to core systems. Companies approaching it with that expectation will encounter technical dependencies, variable model performance, and substantial operating work.

Our conclusion was nevertheless more positive than expected. Hermes provides a flexible environment for combining local models, cloud models, files, terminal operations, browsers, reusable skills, and persistent knowledge. That makes it a compelling platform for technically managed pilots.

The main surprise was how little autonomy is required to create value. A restricted agent that researches information, processes files, and produces a reviewable deliverable can already reduce the workload of experienced employees.

For midmarket businesses, the practical approach lies between two extremes. Hermes should not be treated as a simple chatbot, but it should not receive broad authority over production systems at the beginning either.

A hybrid model is more appropriate. Hermes handles variable knowledge work, evaluates situations, and prepares outputs. Binding transactions continue through defined processes, assigned roles, and human approval. Employees retain professional accountability.

Under that model, the AI agent is not an uncontrolled substitute for staff. It becomes a technical assistant with a limited mandate, observable work, and an escalation path. That is where Hermes currently offers its strongest business value.

Which sources support the statistics used in this article?

Sources for statistics

  1. Bitkom e.V.: “Digitization of the Economy: Almost Every Company Is Engaging with AI”
    https://www.bitkom.org/Presse/Presseinformation/Digitalisierung-der-Wirtschaft-Unternehmen-beschaeftigen-sich-mit-KI
  2. McKinsey & Company: “The State of AI: Global Survey 2025”
    https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
  3. Gartner: “Over 40% of Agentic AI Projects Will Be Canceled by End of 2027”
    https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027

Which further reading provides additional guidance?

Further reading

  1. Nous Research: Official Hermes Agent Documentation
    https://hermes-agent.nousresearch.com/docs/
  2. OWASP: Top 10 for Agentic Applications for 2026
    https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/
  3. NIST: AI Agent Standards Initiative
    https://www.nist.gov/artificial-intelligence/ai-agent-standards-initiative

Frequently asked questions

What is an AI agent for business?

An AI agent for business is software that does more than compose answers. It can complete a task through multiple steps and, when authorized, use files, browsers, terminals, knowledge sources, or APIs. The operating model determines its practical value: permissions, review gates, logging, and a bounded use case matter as much as the underlying language model.

Is Hermes Agent suitable for production use?

Hermes Agent can support production work when its scope is limited, technically managed, and protected by approval steps. It is not automatically prepared for unattended operation in critical processes. Companies need their own testing, role model, audit logging, update procedure, and an accountable owner who reviews outputs as well as changes to skills, tools, and configuration.

Can Hermes Agent run fully locally?

Hermes Agent can use local language models and local execution environments, allowing selected data flows to remain on company-controlled hardware. Fully local does not mean every feature works offline. Web research, external APIs, cloud models, and certain integrations still require network access, credentials, and an intentional technical decision about which information may leave the local environment.

What hardware does Hermes Agent require?

Hardware requirements depend mainly on the selected model, context size, and expected response time. Smaller local models run on modern workstations but may become less reliable during long tool sequences. Larger models need more memory and compute capacity. A practical alternative is to keep tools and files local while using a stronger cloud model for demanding reasoning tasks.

Which tasks does Hermes Agent handle well?

In our test, Hermes performed best on work that combined research, file handling, analysis, and several connected steps. Useful examples include market scans, document preparation, structured evaluations, technical checks, and repeatable knowledge work. It is less suitable when every transaction must follow an identical path, every posting must be exact, or an outcome becomes binding without human review.

How is Hermes Agent different from ChatGPT?

ChatGPT is primarily a general AI interface, while Hermes Agent is an execution environment that connects models with tools, workspaces, skills, memory, and multiple providers. The main difference is therefore the operating framework rather than a single model. Hermes can pursue multi-step tasks and edit files, but it also requires more configuration, access management, and technical ownership.

How secure is running Hermes Agent in Docker?

Docker can reduce exposure by executing commands inside an isolated environment. The actual protection still depends on mounted folders, forwarded credentials, network access, and container privileges. A poorly configured mount can expose sensitive host data. Companies should share only required directories, restrict write access, separate test data, and require approval before destructive or externally visible actions.

Does Hermes Agent need ERP, CRM, or document management integrations?

Production integrations are not required for an initial test. Many useful scenarios can be evaluated with exported files, synthetic data, and a separate workspace. ERP, CRM, or document management access should follow only after the use case performs reliably. Service accounts, minimum permissions, logging, error handling, and a fallback path matter more than connecting as many systems as possible.

Can Hermes Agent replace employees?

Hermes Agent does not replace an entire role. It can take over portions of knowledge work by collecting information, drafting documents, editing files, and accelerating standardized preparation. Professional accountability, prioritization, negotiations, and decisions with financial, legal, or employment consequences remain human responsibilities. The strongest business case is usually experienced staff working with an agent, not an unattended substitute.

How should a midmarket company start with Hermes Agent?

Start with one bounded, recurring, and reviewable use case. The task should have representative sample data, a visible amount of manual effort, and an accountable business owner. Define the workspace, permissions, expected output, and acceptance tests before adding production information. Connect operational systems or scheduled automation only after repeated runs produce dependable results and exceptions are understood.