honglyua 12225ce826 !300 Fix: refresh apt lists before installing numactl in CI prepare.sh
Merge pull request !300 from yicuixi/fix/apt-update-before-numactl
2026-09-22 09:04:49 +00:00
2025-03-26 15:06:41 +08:00
2026-09-17 14:11:33 +08:00
2026-09-09 10:50:33 +08:00
2025-07-25 15:45:01 +08:00
2026-09-15 13:50:12 +08:00
2026-09-15 13:50:12 +08:00
2026-09-15 13:50:12 +08:00

English Chinese

DeepSparkInference

Homepage LICENSE Release

DeepSparkInference ModelZoo, as a core project of the DeepSpark open-source community, was officially open-sourced in March 2024. The first release selected 48 inference model examples, covering fields such as computer vision, natural language processing, and speech recognition. More AI domains will be gradually expanded in the future.

The models in DeepSparkInference provide inference examples and guidance documents for running on inference engines IGIE or ixRT self-developed by Iluvatar CoreX. Some models provide evaluation results based on the self-developed GPGPU Zhikai 100.

IGIE (Iluvatar GPU Inference Engine) is a high-performance, highly gene, and end-to-end AI inference engine developed based on the TVM framework. It supports multi-framework model, quantization, graph optimization, multi-operator library support, multi-backend support, and automatic operator tuning, providing an easy-to-deploy, high-throughput, and low-latency complete solution for inference scenarios.

ixRT (Iluvatar CoreX RunTime) is a high-performance inference engine independently developed by Iluvatar CoreX, focusing on maximizing the performance of Iluvatar CoreX's GPGPU and achieving high-performance inference for models in various fields. ixRT supports features such as dynamic shape inference, plugins, and INT8/FP16 inference.

DeepSparkInference will be updated quarterly, and model categories will be gradually enriched, with large model inference to be expanded in the future.

ModelZoo

LLM (Large Language Model)

Model Engine Supported IXUCA SDK
Baichuan2-7B vLLM ✅ 4.3.0
ChatGLM-3-6B vLLM ✅ 4.3.0
ChatGLM-3-6B-32K vLLM ✅ 4.3.0
CosyVoice2-0.5B PyTorch ✅ 4.3.0
CosyVoice2-0.5B ixRT ✅ 5.0.0
CosyVoice2-0.5B IGIE ✅ 5.0.0
DeepSeek-R1-Distill-Llama-8B vLLM ✅ 4.3.0
DeepSeek-R1-Distill-Llama-70B vLLM ✅ 4.3.0
DeepSeek-R1-Distill-Qwen-1.5B vLLM ✅ 4.3.0
DeepSeek-R1-Distill-Qwen-7B vLLM ✅ 4.4.0
DeepSeek-R1-Distill-Qwen-14B vLLM ✅ 4.3.0
DeepSeek-R1-Distill-Qwen-32B vLLM ✅ 4.3.0
DeepSeek-OCR Transformers ✅ 4.3.0
DeepSeek-OCR vLLM ✅ 5.0.0
ERNIE-4.5-21B-A3B FastDeploy ✅ 4.3.0
ERNIE-4.5-300B-A47B FastDeploy ✅ 4.3.0
ERNIE-4.5-VL-28B-A3B-Thinking Transformers ✅ 4.4.0
GLM-4V vLLM ✅ 4.3.0
InternLM3 LMDeploy ✅ 4.3.0
InternLM3 vLLM ✅ 4.4.0
IndexTTS-2 IndexTTS ✅ 4.4.0
Llama2-7B vLLM ✅ 4.3.0
Llama2-7B TRT-LLM ✅ 4.3.0
Llama2-13B TRT-LLM ✅ 4.3.0
Llama2-70B TRT-LLM ✅ 4.3.0
Llama3-70B vLLM ✅ 4.3.0
E5-V vLLM ✅ 4.3.0
MiniCPM-o-2 vLLM ✅ 4.3.0
MiniCPM-V-2 vLLM ✅ 4.3.0
MiniCPM-V-4 vLLM ✅ 5.0.0
NVLM vLLM ✅ 4.3.0
Phi3_v vLLM ✅ 4.3.0
PaliGemma vLLM ✅ 4.3.0
PaddleOCR-VL Transformers ✅ 4.4.0
Qwen-7B vLLM ✅ 4.3.0
Qwen-VL vLLM ✅ 4.3.0
Qwen2-VL vLLM ✅ 4.3.0
Qwen2.5-VL vLLM ✅ 4.4.0
Qwen1.5-7B vLLM ✅ 4.3.0
Qwen1.5-7B TGI ✅ 4.3.0
Qwen1.5-14B vLLM ✅ 4.3.0
Qwen1.5-32B Chat vLLM ✅ 4.3.0
Qwen1.5-72B vLLM ✅ 4.3.0
Qwen2-7B Instruct vLLM ✅ 4.3.0
Qwen2-72B Instruct vLLM ✅ 4.3.0
Qwen3_Moe vLLM ✅ 5.0.0
Qwen3-8B vLLM ✅ 4.4.0
Qwen3-32B vLLM ✅ 4.4.0
Qwen3-30B-A3B-Thinking vLLM ✅ 4.4.0
Qwen3-235B-A22B-Thinking vLLM ✅ 4.4.0
Qwen3-Next-80B-A3B vLLM ✅ 4.4.0
Qwen3-Embedding-8B vLLM ✅ 4.4.0
Qwen3-ASR-1.7B Qwen-ASR ✅ 4.4.0
Qwen3-TTS-12Hz-1.7B-Base Qwen-TTS ✅ 4.4.0
Qwen3.5-27B vLLM ✅ 5.0.0
Qwen3.6-27B vLLM ✅ 5.0.0
DeepSeek-V3.1 vLLM ✅ 4.4.0
StableLM2-1.6B vLLM ✅ 4.3.0
Step3 vLLM ✅ 4.4.0
Ultravox vLLM ✅ 4.3.0
Whisper vLLM ✅ 4.3.0
XLMRoberta vLLM ✅ 4.3.0

Computer Vision

Classification

Model Prec. IGIE ixRT IXUCA SDK
AlexNet FP16 ✅ ✅ 4.3.0
INT8 ✅ ✅ 4.3.0
CLIP FP16 ✅ ✅ 4.3.0
Conformer-B FP16 ✅ 4.3.0
ConvNeXt-Base FP16 ✅ ✅ 4.3.0
ConvNeXt-Large FP16 ✅ 4.4.0
ConvNext-S FP16 ✅ 4.3.0
ConvNeXt-Small FP16 ✅ ✅ 4.3.0
ConvNeXt-Tiny FP16 ✅ 4.3.0
CSPDarkNet53 FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
CSPResNet50 FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
CSPResNeXt50 FP16 ✅ ✅ 4.3.0
DeiT-B FP16 ✅ 4.4.0
DeiT-tiny FP16 ✅ ✅ 4.3.0
DenseNet121 FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.4.0
DenseNet161 FP16 ✅ ✅ 4.3.0
DenseNet169 FP16 ✅ ✅ 4.3.0
DenseNet201 FP16 ✅ ✅ 4.3.0
DINOv2 FP16 ✅ ✅ 5.0.0
DINOv3 FP16 ✅ 5.0.0
EfficientNet-B0 FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
EfficientNet-B1 FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
EfficientNet-B2 FP16 ✅ ✅ 4.3.0
EfficientNet-B3 FP16 ✅ ✅ 4.3.0
EfficientNet-B4 FP16 ✅ ✅ 4.3.0
EfficientNet-B5 FP16 ✅ ✅ 4.3.0
EfficientNet-B6 FP16 ✅ 4.3.0
EfficientNet-B7 FP16 ✅ 4.3.0
EfficientNetV2 FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
EfficientNetv2_rw_t FP16 ✅ ✅ 4.3.0
EfficientNetv2_s FP16 ✅ ✅ 4.3.0
GoogLeNet FP16 ✅ ✅ 4.3.0
INT8 ✅ ✅ 4.3.0
HRNet-W18 FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
InceptionV3 FP16 ✅ ✅ 4.3.0
INT8 ✅ ✅ 4.3.0
Inception-ResNet-V2 FP16 ✅ 4.3.0
INT8 ✅ 4.3.0
Mixer_B FP16 ✅ 4.3.0
MNASNet0_5 FP16 ✅ 4.3.0
MNASNet0_75 FP16 ✅ 4.3.0
MNASNet1_0 FP16 ✅ 4.3.0
MNASNet1_3 FP16 ✅ 4.3.0
MobileNetV1 FP16 ✅ 4.4.0
MobileNetV2 FP16 ✅ ✅ 4.3.0
INT8 ✅ ✅ 4.3.0
MobileNetV3_Large FP16 ✅ 4.3.0
MobileNetV3_Small FP16 ✅ ✅ 4.3.0
Mobilevit_s FP16 ✅ 4.4.0
MViTv2_base FP16 ✅ 5.0.0
RegNet_x_16gf FP16 ✅ 4.3.0
RegNet_x_1_6gf FP16 ✅ 4.3.0
RegNet_x_3_2gf FP16 ✅ 4.3.0
RegNet_x_8gf FP16 ✅ 4.3.0
RegNet_y_8gf FP16 ✅ 4.4.0
RegNet_x_32gf FP16 ✅ 4.3.0
RegNet_x_400mf FP16 ✅ 4.3.0
RegNet_x_800mf FP16 ✅ 4.3.0
RegNet_y_1_6gf FP16 ✅ 4.3.0
RegNet_y_16gf FP16 ✅ 4.3.0
RegNet_y_3_2gf FP16 ✅ 4.3.0
RegNet_y_32gf FP16 ✅ 4.3.0
RegNet_y_400mf FP16 ✅ 4.3.0
RegNet_y_800mf FP16 ✅ 4.4.0
RepVGG FP16 ✅ ✅ 4.3.0
Res2Net50 FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
ResNeSt50 FP16 ✅ 4.3.0
ResNet101 FP16 ✅ ✅ 4.3.0
INT8 ✅ ✅ 4.3.0
ResNet152 FP16 ✅ 4.3.0
INT8 ✅ 4.3.0
ResNet18 FP16 ✅ ✅ 4.3.0
INT8 ✅ ✅ 4.3.0
ResNet34 FP16 ✅ ✅ 5.0.0
INT8 ✅ 4.3.0
ResNet50 FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
ResNetV1D50 FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
ResNeXt50_32x4d FP16 ✅ ✅ 4.3.0
ResNeXt101_64x4d FP16 ✅ ✅ 4.3.0
ResNeXt101_32x8d FP16 ✅ ✅ 4.3.0
SEResNet50 FP16 ✅ 4.3.0
ShuffleNetV1 FP16 ✅ ✅ 4.4.0
ShuffleNetV2_x0_5 FP16 ✅ ✅ 4.3.0
ShuffleNetV2_x1_0 FP16 ✅ ✅ 4.3.0
ShuffleNetV2_x1_5 FP16 ✅ ✅ 4.3.0
ShuffleNetV2_x2_0 FP16 ✅ ✅ 4.3.0
SqueezeNet 1.0 FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
SqueezeNet 1.1 FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
SVT Base FP16 ✅ 4.3.0
Swin Transformer FP16 ✅ ✅ 4.3.0
Swin Transformer Large FP16 ✅ 4.3.0
Swin-S FP16 ✅ 5.0.0
Twins_PCPVT FP16 ✅ 4.3.0
VAN_B0 FP16 ✅ 4.3.0
VGG11 FP16 ✅ 4.3.0
VGG13 FP16 ✅ 4.3.0
VGG13_BN FP16 ✅ 4.3.0
VGG16 FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
VGG16-BN FP16 ✅ 5.0.0
VGG19 FP16 ✅ 4.3.0
VGG19_BN FP16 ✅ 4.3.0
ViT FP16 ✅ ✅ 4.3.0
ViT-B-32 FP16 ✅ 4.4.0
ViT-L-14 FP16 ✅ 4.4.0
Wide ResNet50 FP16 ✅ ✅ 4.3.0
INT8 ✅ ✅ 4.3.0
Wide ResNet101 FP16 ✅ 4.3.0
YOLOv8n-cls FP16 ✅ 5.0.0

Object Detection

Model Prec. IGIE ixRT IXUCA SDK
ATSS FP16 ✅ ✅ 4.3.0
CenterNet FP16 ✅ ✅ 4.3.0
DETR FP16 ✅ 4.3.0
FCOS FP16 ✅ ✅ 4.3.0
FoveaBox FP16 ✅ ✅ 4.3.0
FSAF FP16 ✅ ✅ 4.3.0
GFL FP16 ✅ 4.3.0
Grounding DINO FP16 ✅ 5.0.0
HRNet FP16 ✅ ✅ 4.3.0
PAA FP16 ✅ ✅ 4.3.0
RetinaFace FP16 ✅ ✅ 4.3.0
RetinaNet FP16 ✅ ✅ 4.3.0
RTMDet FP16 ✅ 4.3.0
RTDETR FP16 ✅ ✅ 5.0.0
INT8 ✅ 5.0.0
SABL FP16 ✅ 4.3.0
SSD FP16 ✅ 4.3.0
YOLOF FP16 ✅ ✅ 4.3.0
YOLOv3 FP16 ✅ ✅ 4.3.0
INT8 ✅ ✅ 4.3.0
YOLOv4 FP16 ✅ ✅ 4.3.0
INT8 ✅ ✅ 4.3.0
YOLOv5m FP16 ✅ ✅ 4.3.0
INT8 ✅ ✅ 4.3.0
YOLOv5s FP16 ✅ ✅ 4.3.0
INT8 ✅ ✅ 4.3.0
YOLOv6s FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
YOLOv7 FP16 ✅ ✅ 4.3.0
INT8 ✅ ✅ 4.3.0
YOLOv8l FP16 ✅ 4.4.0
YOLOv8n FP16 ✅ ✅ 4.3.0
INT8 ✅ ✅ 4.3.0
YOLOv8s FP16 ✅ 4.3.0
INT8 ✅ 4.3.0
YOLOv8x FP16 ✅ 4.4.0
YOLOv9s FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
YOLOv10s FP16 ✅ ✅ 4.3.0
YOLOv10x FP16 ✅ 5.0.0
YOLOv11l FP16 ✅ 4.4.0
YOLOv11m FP16 ✅ 4.4.0
INT8 ✅ 4.4.0
YOLOv11n FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
YOLOv11s FP16 ✅ 4.4.0
INT8 ✅ 4.4.0
YOLOv11x FP16 ✅ 4.4.0
YOLOv12n FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
YOLOv13n FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
YOLOv26n FP16 ✅ 4.4.0
YOLOXm FP16 ✅ ✅ 4.3.0
INT8 ✅ ✅ 4.3.0
Model Prec. PaddlePaddle IXUCA SDK
RTDETR FP16 ✅ 5.0.0
Model Prec. Pytorch IXUCA SDK
YOLOv8n FP16 ✅ 5.0.0

Face Recognition

Model Prec. IGIE ixRT IXUCA SDK
Arcface FP16 ✅ 5.0.0
FaceNet FP16 ✅ 4.3.0
INT8 ✅ 4.3.0
YOLOv8n-Face FP16 ✅ 5.0.0

OCR (Optical Character Recognition)

Model Prec. IGIE ixRT IXUCA SDK
CRNN FP16 ✅ 4.4.0
DBNet FP16 ✅ 4.4.0
Kie_layoutXLM FP16 ✅ 4.3.0
SVTR FP16 ✅ 4.3.0

Pose Estimation

Model Prec. IGIE ixRT IXUCA SDK
HRNetPose FP16 ✅ 4.3.0
Lightweight OpenPose FP16 ✅ 4.3.0
RTMPose FP16 ✅ ✅ 4.3.0
YOLOv8n-pose FP16 ✅ 5.0.0

Instance Segmentation

Model Prec. IGIE ixRT IXUCA SDK
Mask R-CNN FP16 ✅ 4.2.0
SOLOv1 FP16 ✅ 4.3.0

Semantic Segmentation

Model Prec. IGIE ixRT IXUCA SDK
DDRNet FP16 ✅ 4.4.0
UNet FP16 ✅ ✅ 4.3.0

Multi-Object Tracking

Model Prec. IGIE ixRT IXUCA SDK
FastReID FP16 ✅ ✅ 4.3.0
DeepSort FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
RepNet-Vehicle-ReID FP16 ✅ ✅ 4.3.0

Multimodal

Model Engine Supported IXUCA SDK
Aria vLLM ✅ 4.3.0
Chameleon-7B vLLM ✅ 4.3.0
CLIP IxFormer ✅ 4.3.0
DeepSeek-VL2-tiny vLLM ✅ 4.4.0
Fuyu-8B vLLM ✅ 4.3.0
FLUX.1-Dev xDiT ✅ 4.4.0
H2OVL Mississippi vLLM ✅ 4.3.0
HunyuanVideo xDiT ✅ 4.4.0
HunyuanDiT-v1.2 xDiT ✅ 4.4.0
Idefics3 vLLM ✅ 4.3.0
InternVL2-4B vLLM ✅ 4.3.0
LLaVA vLLM ✅ 4.3.0
LLaVA-Next-Video-7B vLLM ✅ 4.3.0
Llama-3.2 vLLM ✅ 4.3.0
Pixtral vLLM ✅ 4.3.0
Qwen-Image ComfyUI ✅ 4.4.0
Stable Diffusion 1.5 Diffusers ✅ 4.3.0
Stable Diffusion 2.1 ixRT ✅ 4.4.0
Stable Diffusion 3 Diffusers ✅ 5.0.0
SD3-Medium xDiT ✅ 4.4.0
Wan2.1-T2V-14B xDiT ✅ 4.4.0
Wan2.2-TI2V-5B xDiT ✅ 4.4.0
Z-Image Diffusers ✅ 4.4.0

NLP

PLM (Pre-trained Language Model)

Model Prec. IGIE ixRT IXUCA SDK
ALBERT FP16 ✅ 4.3.0
BERT Base NER INT8 ✅ 4.3.0
BERT Base SQuAD FP16 ✅ ✅ 4.3.0
INT8 ✅ 4.3.0
BERT Large SQuAD FP16 ✅ ✅ 4.3.0
INT8 ✅ ✅ 4.3.0
DeBERTa FP16 ✅ 4.3.0
RoBERTa FP16 ✅ 4.3.0
RoFormer FP16 ✅ 4.3.0
VideoBERT FP16 ✅ 4.2.0

Audio

Speech Recognition

Model Prec. IGIE ixRT IXUCA SDK
Conformer FP16 ✅ ✅ 4.3.0
DeepSpeech2 FP16 ✅ 4.4.0
Transformer ASR FP16 ✅ 4.2.0

Others

Recommendation Systems

Model Prec. IGIE ixRT IXUCA SDK
Wide & Deep FP16 ✅ 4.3.0

SDK && Docker

You can visit the Iluvatar Developer to obtain the IXUCA software stack and docker container.

Community

Code of Conduct

Please refer to DeepSpark Code of Conduct on Gitee or on GitHub.

Contact

Please contact developers@iluvatar.com.

Contribution

Please refer to the DeepSparkInference Contributing Guidelines.

Disclaimers

DeepSparkInference only provides download and preprocessing scripts for public datasets. These datasets do not belong to DeepSparkInference, and DeepSparkInference is not responsible for their quality or maintenance. Please ensure that you have the necessary usage licenses for these datasets. Models trained based on these datasets can only be used for non-commercial research and education purposes.

To dataset owners:

If you do not want your dataset to be published on DeepSparkInference or wish to update the dataset that belongs to you on DeepSparkInference, please submit an issue on Gitee or Github. We will delete or update it according to your issue. We sincerely appreciate your support and contributions to our community.

License

This project is released under Apache-2.0 License.

S
Description
No description provided
Readme Apache-2.0 50 MiB
Languages
Python 76.9%
Shell 20.9%
C++ 1.4%
Cuda 0.6%