Private AIPrototype · demo on request

Your documents answer.
Nothing leaves the building.

A ChatGPT-class assistant over your contracts, policies, drawings, spreadsheets and scans, running on one GPU box inside your network or in your own EU tenant. Every answer cites the page it came from. No document, no question and no answer ever reaches a third-party API.

0 bytesto external APIs, by architecture
1 boxDGX Spark class, sits under a desk
Open weightsQwen or Mistral, Apache-2.0 stack
Citedevery answer links to page and paragraph
knowledge base
● PRIVATEAccess follows your groups. Legal does not see HR, the shop floor does not see contracts.
Ask anything about this knowledge base…local GPU
source · click a citation
modelQwen · local
retrieved–
latency–
sent outside0 B

Illustrative conversation on fictional documents. The real assistant runs on your files in a two-week pilot.

What people ask it all day.

A private assistant is not a chatbot on the website. It is the colleague who has read everything and always shows where the answer came from.

§

Contract archive

"Which agreements auto-renew this quarter?" "Where do we accept unlimited liability?" Terms, deadlines, risks and differences across your contract archive, with the clause on screen.

in PDF, DOCX, scans → out answer + clause, export to spreadsheet
HR

Policies & internal rules

Leave, travel, procurement, safety. New employees ask the assistant instead of the three people who know. Answers stay consistent with the current version of the policy.

in DOCX, intranet pages → out answer + policy section
⚙

Manuals, drawings, maintenance logs

Thousands of scanned pages in two languages. "Torque for the M20 bolts on the kiln drive?" answered from the manual, with the page shown to the technician.

in scans, PDFs, photos → out answer + page image
▦

Spreadsheets in plain language

"Sales to Print Co in Q2 by product." The assistant reads the tables you already keep and answers with the numbers and the sheet they came from.

in XLSX, CSV → out figures + sheet reference
✉

Drafting with your own material

Reply to a tender, summarise a 90-page report, draft a clause in your house style. The model works from your documents, not from the internet.

in your archive → out draft with sources
🎙

Voice and field notes

Dictated inspection notes and meeting recordings transcribed locally and filed into the same searchable base.

in audio → out text, searchable, cited

The demo corpus for a call is small: two or three scenarios on your own files, an Excel table, a contract set, a scanned archive.

Book a demo

What stays inside, and why it is not a toggle.

Cloud assistants promise not to train on your data. Ours cannot see it. The whole stack, from the model weights to the vector index, runs on hardware you own or in a tenant you control.

  • 1
    Open-weight models you can inspect. Qwen or Mistral class, Apache-2.0 licensed, pinned to a version. No silent model swaps, no usage-based pricing.
  • 2
    Retrieval with citations, not memory. The model answers from the passages retrieved for that question and must show them. Wrong answers are visible, and correctable.
  • 3
    Your groups, your rights. Knowledge bases inherit access from your directory. The assistant cannot leak a document to someone who could not open it anyway.
  • 4
    Monitoring you can read. Latency, load, failed queries and who asked what, in Grafana on the same box. Useful for the EU AI Act file.
Qwen / MistralvLLMQdrantBGE-M3 + rerankerWhisperFastAPInginx · TLSPrometheus · GrafanaDocker Compose
Your documentsDMS · shares · email
Your usersSSO · groups
↓ index once, then only deltas ↓
One GPU box · DGX Spark classLLM · embeddings · vector DB · API · UI
↓ answers with page references ↓
Web assistantbrowser, no install
IntegrationsCRM · ERP · kiosk · Teams
Traffic to the internet from this system0 B

How a pilot runs.

Four steps, about six weeks, on your documents. You judge the answers, not our slides.

01 · week 1

Discovery

Which knowledge bases, who may see what, where the documents live, which systems to connect. Hardware sizing and ordering.

02 · weeks 2–3

Demo corpus

Your real files: contracts, Excel, scans, voice material. Indexed on the box, on your premises or in your EU tenant.

03 · weeks 3–5

Evaluation

Your people ask their real questions. We measure answer quality, citation correctness and speed together, and write it down.

04 · week 6 →

Rollout

Integrations, access groups, training for the team, monitoring handover. Support and model updates on a monthly basis.

Where we are honest before you ask.

This direction is a working prototype, not a product with a hundred installations. Here is what that means.

It is a prototype.

The stack runs end to end on our own GPU and has been evaluated on a real multi-hundred-page corpus. It has not yet run for months inside a client's network. We show it live and price the pilot accordingly.

You need a GPU.

One DGX Spark class machine, or an equivalent server, bought by you and owned by you. We specify it, you keep it. Without it the private guarantee does not exist.

Quality depends on your documents.

Clean PDFs answer well. Faded scans and handwritten notes go through the Document AI pipeline first. We measure on your corpus in the pilot before promising anything.

It will not know what is not written down.

The assistant answers from your archive. If the answer lives in someone's head, it says so instead of inventing one. That is the point of citations.

Asked on every first call

How is this different from ChatGPT Enterprise or Copilot?+

Those run in someone else's cloud under a contract that says they will not misuse your data. This runs on your hardware, so the question does not arise. You also choose and pin the model, and there is no per-seat or per-token bill.

Can it run in our cloud instead of on-premise?+

Yes, in your own EU tenant on a GPU instance you control. The architecture is identical; the difference is who owns the metal.

Does it hallucinate?+

Any language model can. This one must cite the passage it used, and it is told to say "not in the documents" when retrieval finds nothing relevant. Wrong answers become visible and are corrected in the pilot evaluation, not discovered in production.

What about Bulgarian, German, Russian?+

The models and embeddings are multilingual; the corpus we evaluated on was not in English. Your languages are part of the pilot evaluation.

What does it cost?+

The GPU box is a one-off purchase you own. The pilot is fixed-scope. After that, a monthly fee for support and model updates, no per-token charges.

Bring the questions your team asks every week.

Thirty minutes. We show the assistant live, discuss which documents would go in first, and tell you honestly whether a private setup pays off for your size.

Book a call with Artem