Skip to content

Guides ·

What hardware do you need to run AI in-house?

Running AI in-house needs a server with one or more graphics processing units (GPUs), plenty of fast memory, and enough storage for your company’s history. How much you need depends on the number of people, the kind of work, how much history you bring in and how many people use it at once. Seesa’s pricing page shows one example, a dual RTX PRO 6000 server, but the right hardware for your company is confirmed after a discussion, and you can buy it, lease it or use suitable hardware you already own.

What the parts do

An AI model is, in practice, a very large set of numbers that has to be held in memory and worked through for every request. Three parts of a server matter most.

  • GPUs do the heavy calculation. They are far faster than ordinary processors at the kind of arithmetic AI models need, and the model sits in the GPU’s own memory while it works.
  • System memory supports everything else that is running alongside the models.
  • Storage holds your company’s records and the searchable index built from them, so it grows with the amount of history you bring in.

What decides the size

Seesa’s on-premise page says people, history, simultaneous use and the work Seesa will do all affect the installation, and that storage and processing capacity are assessed together before hardware is recommended. The pricing page uses the same inputs: the number of people, the main workload and the history to bring in.

The workload matters because asking questions and drafting is lighter than regular agent work, and heavy agents and code are heavier again. The history matters because more history means more to store and more to index. Simultaneous use matters because several people asking at the same moment need capacity at the same moment.

An example from the pricing page

Seesa’s pricing page shows an indicative configuration: a dual RTX PRO 6000 server, indicative for five simultaneous active requests. The page describes it as two separate 96 GB GPU memory domains, with chat sharing a card with embeddings, a 512 GB ECC memory target and a target of at least 4 TB of usable SSD, with a supplier quote needed. Its planning allowance is one chat replica, with up to eight active requests each.

For five years of history, the page estimates 0.38 TB of source data and 0.12 TB of raw embeddings, with a primary working-storage allowance of 1.0 TB before backup or replication. It adds that additional storage may be needed and is not included in that hardware price.

Two cautions apply. The page states plainly that capacity has not been benchmarked on this exact hardware, and that Seesa verifies mixed-model latency and concurrency before proposing hardware. It is a planning example, not a specification for your company.

Three ways to get the hardware

You do not have to buy a server to use Seesa.

  • Buy it. Pentatonic supplies and installs hardware sized for your team. You own the machine, and Pentatonic manages the Seesa installation.
  • Lease it. Spread the hardware cost over 24, 36 or 48 months. Pentatonic supplies, installs and manages it, and your proposal sets out the lease and software costs separately.
  • Use your own. Share the specifications of suitable hardware you already have, and Pentatonic checks it against your workload and agrees the installation and support arrangements.

What your premises need to provide

The practical side is modest. Your team provides a suitable location, power and network access for the agreed hardware, people who can authorise the system connections you want to use, and someone who can help agree permissions and check the first workflows. Pentatonic then manages configuration, software updates and maintenance for the systems it supplies.

Prices change, so this guide does not quote them. The pricing page has an estimator that shows how people, workload and history move an indicative figure, and the final price is confirmed with you.

How to get an answer for your company

The quickest route is a conversation. Tell Seesa how many people will use it, what you want it to do, which systems it will connect to and how much history it needs to hold, and the hardware recommendation follows from that. Seesa serves companies of all sizes, and the on-premise page sets out the setup process from assessment through to putting it to work.

Why history changes the quote

It is easy to think only about the models, but your company’s records need somewhere to live. Seesa’s pricing page notes that the history you bring in informs the import and storage plan, and that actual source volume may change the hardware quote.

That is one reason the final specification is confirmed with you rather than read from a table. The pricing page also says that standard installation, maintenance and first-line support are included in the standard offer, while bespoke work is scoped separately.

Questions, answered.

Do we have to buy a server?

No. You can buy hardware, lease it over 24, 36 or 48 months, or ask Pentatonic to assess suitable hardware you already have.

How much hardware does our company need?

It depends on the number of people, the workload, the amount of history and how many people use Seesa at the same time. Pentatonic assesses storage and processing capacity together before recommending hardware.

Has the example server been benchmarked?

Seesa’s pricing page says capacity has not been benchmarked on that exact hardware, and that mixed-model latency and concurrency are verified before hardware is proposed.

Talk to us about your company.

Tell us about your team, your systems and what you would like Seesa to take on.

Talk to us