Building a Small Project: Ollama + Neon + Next.js
The Stack
- Ollama |-- Local LLM inference (no API keys, runs on your machine)
- Neon |-- Serverless PostgreSQL (free tier, managed backups)
- Next.js |-- Full-stack React framework (API routes + frontend in one deploy)
Why This Combo?
| Feature | Ollama | Neon | Next.js | | --- | --- | --- | --- | | Cost | $0 (local) | Free tier | $15/mo / Free (Vercel) | | Setup Time | 5 min | 3 min | Already have it | | Scaling | CPU-limited | Auto-scale | Serverless | | Lock-in | None (open source) | Easy exit (PostgreSQL standard) | None (Node-based) |
Why not use OpenAI? API costs add up fast ($0.01-0.10 per request). Ollama lets you experiment free, iterate quickly, keep data private. Neon handles schema and backups so you don't have to manage Postgres yourself.
The Architecture
Browser
||o
Next.js Frontend (app/page.tsx)
||o
Next.js API Route (app/api/generate/route.ts)
||o (fetch)
|-- | Ollama (localhost:11434) for inference
|| Neon PostgreSQL for persistence
||o
Response back to browser
Setup: The Full Stack
1. Install Ollama
Go to ollama.ai/download install for your OS.
# macOS/Linux
wget [ollama.ai/download](https://ollama.ai/download)
chmod +x ollama
./ollama serve
# Windows: Download installer and run
Ollama listens on localhost:11434 by default. Verify:
curl [localhost:11434/api/tags](http://localhost:11434/api/tags)
# Returns: {"models":[]}
2. Pull a Model
ollama pull mistral # 7B, balanced (fastest)
ollama pull llama2 # 7B, good for chat
ollama pull neural-chat # 7B, optimized for conversation
First pull takes 5-10 min (downloads ~5GB). Subsequent runs load from cache.
Gotcha: Don't pull multiple models if you're short on disk/RAM. Start with
mistral|-- it's the fastest and works for most tasks.
3. Create Neon Database
- Sign up at console.neon.tech
- Create project | copy connection string
- Set
DATABASE_URLin.env.local:
DATABASE_URL=postgres://user:password@host/dbname?sslmode=require
4. Schema
Create tables in Neon console or via psql:
CREATE TABLE IF NOT EXISTS prompts (
id SERIAL PRIMARY KEY,
prompt TEXT NOT NULL,
model VARCHAR(50) NOT NULL,
response TEXT,
tokens_used INT,
created_at TIMESTAMPTZ DEFAULT NOW(),
updated_at TIMESTAMPTZ DEFAULT NOW()
);
CREATE INDEX idx_prompts_created ON prompts(created_at DESC);
Building the App
API Route: Generate & Store
// app/api/generate/route.ts
import { Pool } from 'pg';
const pool = new Pool({ connectionString: process.env.DATABASE_URL });
export async function POST(request: Request) {
try {
const { prompt, model = 'mistral' } = await request.json();
if (!prompt) {
return Response.json({ error: 'Prompt required' }, { status: 400 });
}
// Call Ollama
const ollamaRes = await fetch('[localhost:11434/api/generate](http://localhost:11434/api/generate) {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
model,
prompt,
stream: false,
}),
});
if (!ollamaRes.ok) {
throw new Error(`Ollama error: ${ollamaRes.statusText}`);
}
const { response, eval_count } = await ollamaRes.json();
// Store in Neon
await pool.query(
'INSERT INTO prompts (prompt, model, response, tokens_used) VALUES ($1, $2, $3, $4)',
[prompt, model, response, eval_count]
);
return Response.json({ response, model, tokens: eval_count });
} catch (error) {
console.error(error);
return Response.json({ error: 'Failed to generate' }, { status: 500 });
}
}
Frontend: Simple Form
'use client';
import { useState } from 'react';
export default function Generate() {
const [prompt, setPrompt] = useState('');
const [response, setResponse] = useState('');
const [loading, setLoading] = useState(false);
const handleSubmit = async (e: React.FormEvent) => {
e.preventDefault();
setLoading(true);
try {
const res = await fetch('/api/generate', {
method: 'POST',
body: JSON.stringify({ prompt, model: 'mistral' }),
});
const data = await res.json();
setResponse(data.response);
} catch (error) {
console.error(error);
setResponse('Error generating response');
} finally {
setLoading(false);
}
};
return (
<form onSubmit={handleSubmit} className="space-y-4">
<textarea
value={prompt}
onChange={(e) => setPrompt(e.target.value)}
placeholder="Ask Ollama anything..."
className="w-full border rounded p-2"
/>
<button
type="submit"
disabled={loading}
className="px-4 py-2 bg-blue-600 text-white rounded hover:bg-blue-700 disabled:opacity-50"
>
{loading ? 'Generating...' : 'Generate'}
</button>
{response && (
<div className="p-4 bg-gray-100 rounded whitespace-pre-wrap">
{response}
</div>
)}
</form>
);
}
Common Issues & Fixes
"Connection refused: localhost:11434"
Ollama not running. Start it:
./ollama serve
"FATAL: password authentication failed for user"
Wrong DATABASE_URL. Check Neon console | Connection string. Copy the full URL including ?sslmode=require.
"Error: Ollama connection timeout"
Model too slow or stuck. Check Ollama logs:
# If using ollama serve in background, check process
ps aux | grep ollama
Restart if needed:
killall ollama
./ollama serve
Memory issues on large models
Ollama supports quantized models (4-bit, 5-bit) to reduce size:
ollama pull mistral:7b-q4_0 # 4-bit quantized, ~4GB RAM
Deployment Options
Vercel + Neon + Local Ollama
Frontend + API routes | Vercel Database | Neon Ollama | Your machine (webhook triggers it)
This works but means requests timeout if Ollama isn't running. Better for batch jobs.
// app/api/queue/route.ts |-- Queue job to run later
export async function POST(request: Request) {
const job = await request.json();
// Store in Neon, have your local Ollama machine poll for jobs
await pool.query('INSERT INTO jobs (prompt, status) VALUES ($1, $2)', [job.prompt, 'queued']);
return Response.json({ id: job.id, status: 'queued' });
}
Docker + Compose
Run all three locally:
version: '3.9'
services:
ollama:
image: ollama/ollama:latest
ports:
- '11434:11434'
volumes:
- ./ollama-data:/root/.ollama
environment:
- OLLAMA_KEEP_ALIVE=5m
web:
build: .
ports:
- '3000:3000'
environment:
- DATABASE_URL=postgres://user:pass@postgres:5432/db
depends_on:
- postgres
postgres:
image: postgres:16-alpine
environment:
POSTGRES_PASSWORD: secret
volumes:
- ./postgres-data:/var/lib/postgresql/data
Example: vhsbox
vhsbox.netlify.app uses this exact stack to tag video metadata with Ollama, store in Neon, serve from Next.js. See the repo for full implementation.
Performance & Costs
| Scenario | Monthly Cost | Notes | | ---------- | ------------- | ------- | | 1000 requests/mo | $0 Ollama + $0 Neon + $0 Vercel | Free tier covers it all | | 10k requests/mo | $0 Ollama + $0 Neon + $0 Vercel | Still free; Neon free tier is generous | | 100k requests/mo | $0 Ollama + $15 Neon | Neon scales, Ollama is local | | 1M requests/mo | $0 Ollama + $300 Neon | Storage/compute adds up; consider splitting to OpenAI |
When to switch to OpenAI: If you hit 100k+ requests/mo and Neon costs spike, compare against OpenAI API. At that scale, trade-offs flip |-- API costs predictable, your Ollama machine becomes bottleneck.
Takeaway
Ollama + Neon + Next.js is the fastest path to a working AI-powered fullstack app with zero billing. Ship the MVP, measure demand, scale when needed. The pattern survives that transition |-- swap Ollama for OpenAI API, Neon for self-managed Postgres, all with one line changes.
Related Resources
- Ollama Documentation
- Neon PostgreSQL Docs
- Next.js App Router Guide
- Vercel Deployment Docs
- vhsbox Live Demo
Key Takeaways
- Ollama runs LLMs locally with zero API costs and full data privacy
- Neon provides serverless PostgreSQL with generous free tier
- Next.js API routes eliminate need for separate backend service
- Local-first development means faster iteration and offline work
- Cost stays at $0 until you need production scale
Further Reading
- Official Next.js docs: Next.js Docs
- Docker best practices: Docker Best Practices
- Ollama guide: Ollama Guide
Ollama Model Pull
ollama pull mistralThanks for reading!
Read More Articles