Skip to content

Isvik.cpp

Isvik.cpp logo on a dark banner

Original Isvik logo
Original Isvik logo · © 2026 AbyssGG

Built for OpenVINO and local AI on Intel AI PCs.
C++20 · OpenVINO GenAI · Optional ONNX support · Desktop chat · Interactive CLI · Local APIs

Isvik.cpp source version 0.1.0 OpenVINO GenAI 2026.4.0.0 Optional ONNX Runtime GenAI 0.17.0 C++20 Experimental NVIDIA TensorRT GGUF support Apache 2.0 license Build status

English | 简体中文

What is Isvik.cpp?

Isvik.cpp is a C++20 local AI runtime designed specifically to make OpenVINO language-model inference accessible on Intel AI PCs. It provides a desktop chat interface, a persistent bilingual CLI, a model library, saved memories, and a local API server. Inference runs on your hardware without a cloud account. Compatible ONNX Runtime GenAI model packages are supported through an optional backend; NVIDIA TensorRT is an experimental integration.

Built for OpenVINO on Intel AI PCs

The project began with a practical goal: make OpenVINO language-model inference easy to use from a desktop app and a real terminal, then expose the same local model through an API. OpenVINO GenAI is the primary path for OpenVINO IR models and the GGUF models supported by the installed runtime. You can select CPU, Intel GPU, or NPU when the model and device support that combination.

The model library, chat, CLI, and API share the same runtime services. This lets you inspect a model, choose a device, and use it in the interface that fits your workflow. See model support for the exact compatibility limits.

Supported backends

Backend Role Model input Where it runs
OpenVINO GenAI Primary runtime OpenVINO IR and selected GGUF models Supported CPU, GPU, or NPU devices
ONNX Runtime GenAI Optional Compatible ONNX text-generation packages Providers available in the build
NVIDIA TensorRT Experimental Supported Gemma 4 GGUF models through Isvik's native plugin NVIDIA GPU
TensorRT-RTX provider Optional Compatible ONNX Runtime GenAI packages Supported NVIDIA GPU, with configured fallback

Model recognition does not guarantee inference support. Architecture, tensor encoding, tokenizer assets, installed SDKs, and the selected device all matter. The TensorRT GGUF implementation has limited model coverage and known performance costs; it is still being developed. See model support.

Versions

Isvik.cpp is at source version 0.1.0, as set in CMake and the Windows executable resource. The Windows x64 beta package is published as isvik-b010-x64.

Component Version used by this repository
OpenVINO GenAI and its bundled runtime (Windows) 2026.4.0.0
ONNX Runtime GenAI (Windows) 0.17.0
ONNX Runtime (Windows) 1.26.0
Slint C++ 1.18.1
TensorRT-RTX execution provider ABI (optional, Windows) 0.4.2
TensorRT SDK and CUDA Toolkit (optional) Versions from the local installation; no fixed patch version

The pinned package values come from the CMake integration files. A backend is included only when its required SDK is available in the selected build.

Get started

Build on Windows

Install Visual Studio 2026 with C++ support, CMake 3.21 or later, and Git. From the repository root, run:

cmake --preset windows-msvc-release
cmake --build --preset windows-msvc-release --target Isvik --parallel
./out/build/windows-msvc-release/Release/Isvik.exe

The project is named Isvik.cpp; the executable is named Isvik.exe. The Windows preset enables optional backends when their SDKs are available. See the build guide for other configurations.

Build on Linux

Linux support has not been validated. The following build commands are provisional; the desktop app, OpenVINO inference, and optional backends have not been confirmed to work on Linux.

Use a C++20 GCC toolchain and Ninja:

cmake --preset linux-gcc-release
cmake --build --preset linux-gcc-release --parallel

The supplied Linux preset disables OpenVINO by default. Configure optional backends with their SDKs before using them.

Use the interactive CLI

On Windows, start the terminal workbench directly:

./out/build/windows-msvc-release/Release/Isvik.exe -cli

The prompt stays open for chat and model management. Enter -ls to list imported models, a model number to select one, and -params to inspect the active settings. Each completed reply reports input tokens, output tokens, elapsed time, and tokens per second. Use -lang zh or -lang en to switch language, and -exit to quit. Run Isvik.exe --help for all commands.

To run a single prompt:

./out/build/windows-msvc-release/Release/Isvik.exe --run --model "<path-to-model>" --device CPU --prompt "Explain local inference." --max-tokens 64

Choose a backend and device supported by your model. The model support guide explains the available combinations.

Start the local API

./out/build/windows-msvc-release/Release/Isvik.exe --server --model "<path-to-model>" --backend openvino

The default address is http://127.0.0.1:1234. The server exposes Isvik-native endpoints, OpenAI Chat Completions, and Anthropic Messages. See the API guide for endpoints, streaming, and authentication.

How TensorRT runs GGUF

The experimental native TensorRT path reads a supported Gemma 4 GGUF file in place. Isvik parses its metadata, uses its embedded BPE tokenizer, and reads packed tensor blocks from the original file. A custom TensorRT IPluginV3 passes those packed bytes to CUDA matrix-vector kernels. Isvik's C++ runtime handles the remaining decoding steps, including attention and the KV cache.

This path does not require a GGUF-to-ONNX conversion. TensorRT engines for the custom operation are cached separately; the GGUF model file is not rewritten. Current support is limited to Gemma 4 text models with supported tensor encodings, and this implementation can be slow because it repeatedly reads weights and moves data between host and GPU.

Read the TensorRT GGUF source guide for the exact source files, data flow, supported encodings, and a run command.

Documentation

Guide Description
Build Isvik.cpp Requirements, presets, and optional backends
Model support Formats, backends, and limitations
TensorRT GGUF source guide How the native GGUF path works in source code
API guide Isvik, OpenAI, and Anthropic endpoints
Architecture Application, core, and backend layers
Contributing How to contribute
Documentation style Writing and C++ style conventions

See the code of conduct, security policy, and third-party notices.

License

Isvik.cpp is licensed under the Apache License 2.0. Third-party dependencies retain their own license terms.

The Isvik.cpp logo and wordmark carry a © 2026 AbyssGG notice. The Apache-2.0 license does not grant trademark rights to the project name or logo. See the branding guidance before reusing them.

About

Built for OpenVINO and Intel AI PCs: local AI desktop app, interactive CLI, and API server with optional ONNX support and experimental TensorRT GGUF.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages