Pearl Llama 3 8B Instruct Lite
Cost-efficient 8B instruction model at 80% discounted pricing
About model
Pearl Llama 3 8B Instruct Lite is Meta's 8B parameter instruction-tuned model running on Pearl's Proof of Useful Work protocol, delivering identical performance at 80% discounted pricing. Optimized with Together Lite INT4 quantization for maximum cost efficiency, the model is built for high-volume production workloads where cost per token is the primary constraint. The Pearl kernel extracts cryptographic mining proofs during standard inference without affecting model quality or throughput.
80%
Powered by Pearl's Proof of Useful Work protocol
$0.02
Input and output at the lowest cost tier
8B
INT4 Lite-optimized for maximum throughput
- 80% Discounted Pricing: Powered by Pearl's Proof of Useful Work — mining rewards subsidize inference costs with zero impact on model quality or throughput
- Maximum Cost Efficiency: Together Lite INT4 quantization at $0.02 per million tokens for high-volume production workloads
- Instruction Following: General-purpose conversational AI, basic coding, and content generation with multilingual support
- High Throughput: Optimized for applications with millions of daily requests where latency and cost efficiency are critical
API usage
Endpoint:
Model card
Architecture Overview:
• 8B parameter dense transformer optimized with Together Lite for maximum cost efficiency via INT4 quantization
• 8K token context window
• Instruction-tuned for conversational AI, basic coding, and general-purpose tasks
Pearl Proof of Useful Work:
• This endpoint runs on Pearl's Proof of Useful Work (PoUW) protocol, which extracts cryptographic mining proofs as a side effect of standard AI inference
• Model quality and throughput are preserved — the Pearl kernel operates at the matrix multiplication level without affecting model outputs
• Mining rewards flow back as a direct subsidy, enabling 80% discounted pricing
• Zero-knowledge proofs ensure no model weights or user data are exposed
Performance Characteristics:
• Identical performance to standard Llama 3 8B Instruct Lite — all benchmarks and capabilities are preserved
• Together Lite INT4 optimization for maximum capacity and lowest cost per token
• Suitable for high-volume, cost-sensitive production workloads
Prompting
Together AI API Access:
• Access Pearl Llama 3 8B Instruct Lite via Together AI APIs using the endpoint meta-llama/pearl-Meta-Llama-3-8B-Instruct-Lite
• Authenticate using your Together AI API key in request headers
• $0.02 per million input tokens / $0.02 per million output tokens (80% discount)
• Available on Together AI serverless infrastructure
Applications & use cases
High-Volume Production:
• Cost-efficient inference for classification, routing, and lightweight generation tasks
• High-capacity serving for applications with millions of daily requests
• Together Lite INT4 optimization for maximum throughput per dollar
Conversational AI:
• General-purpose instruction following for chatbots and assistants
• Quick-response applications where latency matters more than maximum reasoning depth
• Multilingual support for diverse user interactions
Development & Prototyping:
• Rapid prototyping and experimentation at minimal cost
• Basic code generation, explanation, and debugging
• Content generation and summarization for production pipelines
- TypeReasoning
- FeaturesFunction Calling
- DeploymentServerless
- Parameters8B
- Context length8K
- Input price
$0.02 / 1M tokens
- Output price
$0.02 / 1M tokens
- Input modalitiesText
- Output modalitiesText
- Quantization levelINT4
- CategoryChat
