AI2025-04-175 min read

Secure by Design: Why Local AI Matters in Avatar Development

As AI becomes more embedded in learning, training, and recruitment, the infrastructure we choose to build with matters more than ever. Privacy concerns, data control,...

As AI becomes more embedded in learning, training, and recruitment, the infrastructure we choose to build with matters more than ever. Privacy concerns, data control, and performance bottlenecks are now top of mind — especially for public sector organisations, healthcare providers, and universities.

At Pxl-Persona, we take a “Secure by Design” approach. That means we don’t just build amazing avatars — we build systems that protect data, ensure control, and scale securely using local AI infrastructure where needed.

This article breaks down why local-first AI architecture matters, when cloud isn’t good enough, and how Pxl-Persona is shaping the next generation of AI avatars for interview training, education, and healthcare.

Why Cloud-Only AI Isn’t Always the Answer

Most AI startups today build on cloud-only platforms. It's fast, flexible, and lets them scale quickly — but it also comes with trade-offs:

  • Data privacy risks (especially with CVs, job interviews, or mental health data)

  • Vendor lock-in with black-box APIs

  • Latency or downtime from third-party LLM providers

  • Limited customisation over model personality or tuning

This is a major problem for NHS trusts, universities, and regulated sectors — all of whom want control, reliability, and accountability.

Why We Build Local-First

At Pxl-Persona, we run and manage our AI stack in-house. Our avatars are built, tested, and run on secure office servers, using NVIDIA A6000s for AI inferencing, with the option to stream to users via the browser using NVIDIA Graphics Cards for the rendering. The 3D avatars utilise PixelStreaming, adding additional security due to the Avatars connecting to our office machine, meaning access to the AI that processes the conversation is only within our network.

This means:

  • Data processing is conducted on our servers (data is securely encrypted)

  • Avatars can run 24/7 without internet dependency (on a local network)

  • We control LLM selection, behaviour tuning, and updates

  • Clients can choose between local or hybrid cloud deployment

Our Stack: Custom, Secure, Flexible

We build avatars using a deeply integrated stack:

  • Unreal Engine 5 for high-fidelity 3D rendering

  • Metahumans + Blender for character design

  • RAG architecture to inject role-specific knowledge (CVs, industry terms, tone)

  • LangChain/Ollama/LlamaCpp for local model orchestration and chaining

  • Encrypted streaming of 3D avatars to browser using PixelStreaming

  • Custom GPU pipelines for secure AI voice model generation

When required, avatars can also be deployed using NVIDIA’s NIM (NVIDIA Inference Microservices) — enabling containerised, enterprise-grade inference.

Read about our RAG and prompt engineering approach (Coming Soon) Explore our technical breakdown of avatar architecture (Coming Soon)

What is the difference between LlamaIndex / LangChain and Ollama / LlamaCPP?

**LangChain **and **LlamaIndex **act as o_rchestrators_. They manage the connections between **RAG **(Retrieval-Augmented Generation), memory, prompts, and components when interacting with a language model.

On the other hand, **Ollama **and **LlamaCPP **are what interface with the LLM on your local machine — they enable you to run language models locally.

In short:

  • Ollama and LlamaCPP let you run and interact with the LLM locally.

  • LangChain and LlamaIndex let you attach tools and data (like RAG, memory, system prompts) to enhance the LLM's capabilities during a conversation.

Security in Practice: How It Works

Let’s break it down:

  • Avatar Rendering → Generated in Unreal Engine and Blender locally

  • Voice AI Generation → Processed on Scenegraphs machines, deployed locally

  • LLM Conversations → Run on local stacks or NIM instances

  • Encrypted Stream → Browser session only sees the final result — no access to raw data

  • RAG data upload → Data is uploaded to our local servers and accessed internally, improving security. Data examples are CV's, job descriptions, scenarios, company data.

This keeps user input, AI prompts, responses, and system behaviour within a secure environment as well as fast because their are no Cloud APIs. It’s ideal for:

  • Universities, colleges, schools and training providers handling student data, or student conversations

  • NHS staff training applications

  • Employment and HR services handling sensitive histories

Secure Doesn’t Mean Static

A common misconception is that “local” means “limiting”. In our case — it’s the opposite.

Because we control the stack:

  • We custom-tune models to simulate personalities (e.g. cold HR interviewer)

  • We adapt scenarios to different industries (construction, healthcare, games)

  • We build realistic avatars with backstory, bias, and tone — using local memory, not API-based functions

  • Utilising Pxl-Personas servers and architecture, or deploy on your premises

Bipolar simulation

What Secure by Design Really Means

Being Secure by Design isn’t just a tech feature. It’s a design philosophy:

  • Build with privacy in mind

  • Make systems auditable and adjustable

  • Support offline-first, not cloud-dependent workflows

  • Deliver high performance using trusted hardware and open-source models

Our goal isn’t just realism — it’s realism with resilience.

Final Thoughts

If you’re building training solutions or assessment tools using AI, **you need more than a clever chatbot.**You need:

  • A controlled, reliable system

  • Secure, scalable deployment options

  • Realism that reflects your audience — not generic outputs

That’s what Pxl-Persona delivers.

Scenegraph Studios project image

Want to learn more about how our avatars run locally and securely in your environment?Book a demo or explore our technical documentation and whitepapers.

CONTACT us today.

Build Confidence with Every Practice

Repeatable, flexible, and student-led.