Development

Building a Small Project: Ollama + Neon + Next.js

Austin H.•August 20, 2026•14 min read
#ollama#neon#nextjs#ai#fullstack

The Stack

  • Ollama |-- Local LLM inference (no API keys, runs on your machine)
  • Neon |-- Serverless PostgreSQL (free tier, managed backups)
  • Next.js |-- Full-stack React framework (API routes + frontend in one deploy)

Why This Combo?

| Feature | Ollama | Neon | Next.js | | --- | --- | --- | --- | | Cost | $0 (local) | Free tier | $15/mo / Free (Vercel) | | Setup Time | 5 min | 3 min | Already have it | | Scaling | CPU-limited | Auto-scale | Serverless | | Lock-in | None (open source) | Easy exit (PostgreSQL standard) | None (Node-based) |

Why not use OpenAI? API costs add up fast ($0.01-0.10 per request). Ollama lets you experiment free, iterate quickly, keep data private. Neon handles schema and backups so you don't have to manage Postgres yourself.

The Architecture

Browser
  ||o
Next.js Frontend (app/page.tsx)
  ||o
Next.js API Route (app/api/generate/route.ts)
  ||o (fetch)
  |-- |    Ollama (localhost:11434) for inference
  ||    Neon PostgreSQL for persistence
  ||o
Response back to browser

Setup: The Full Stack

1. Install Ollama

Go to ollama.ai/download install for your OS.

# macOS/Linux
wget [ollama.ai/download](https://ollama.ai/download)
chmod +x ollama
./ollama serve

# Windows: Download installer and run

Ollama listens on localhost:11434 by default. Verify:

curl [localhost:11434/api/tags](http://localhost:11434/api/tags)
# Returns: {"models":[]}

2. Pull a Model

ollama pull mistral      # 7B, balanced (fastest)
ollama pull llama2       # 7B, good for chat
ollama pull neural-chat  # 7B, optimized for conversation

First pull takes 5-10 min (downloads ~5GB). Subsequent runs load from cache.

Gotcha: Don't pull multiple models if you're short on disk/RAM. Start with mistral |-- it's the fastest and works for most tasks.

3. Create Neon Database

  1. Sign up at console.neon.tech
  2. Create project | copy connection string
  3. Set DATABASE_URL in .env.local:
DATABASE_URL=postgres://user:password@host/dbname?sslmode=require

4. Schema

Create tables in Neon console or via psql:

CREATE TABLE IF NOT EXISTS prompts (
  id SERIAL PRIMARY KEY,
  prompt TEXT NOT NULL,
  model VARCHAR(50) NOT NULL,
  response TEXT,
  tokens_used INT,
  created_at TIMESTAMPTZ DEFAULT NOW(),
  updated_at TIMESTAMPTZ DEFAULT NOW()
);

CREATE INDEX idx_prompts_created ON prompts(created_at DESC);

Building the App

API Route: Generate & Store

// app/api/generate/route.ts
import { Pool } from 'pg';

const pool = new Pool({ connectionString: process.env.DATABASE_URL });

export async function POST(request: Request) {
  try {
    const { prompt, model = 'mistral' } = await request.json();

    if (!prompt) {
      return Response.json({ error: 'Prompt required' }, { status: 400 });
    }

    // Call Ollama
    const ollamaRes = await fetch('[localhost:11434/api/generate](http://localhost:11434/api/generate) {
      method: 'POST',
      headers: { 'Content-Type': 'application/json' },
      body: JSON.stringify({
        model,
        prompt,
        stream: false,
      }),
    });

    if (!ollamaRes.ok) {
      throw new Error(`Ollama error: ${ollamaRes.statusText}`);
    }

    const { response, eval_count } = await ollamaRes.json();

    // Store in Neon
    await pool.query(
      'INSERT INTO prompts (prompt, model, response, tokens_used) VALUES ($1, $2, $3, $4)',
      [prompt, model, response, eval_count]
    );

    return Response.json({ response, model, tokens: eval_count });
  } catch (error) {
    console.error(error);
    return Response.json({ error: 'Failed to generate' }, { status: 500 });
  }
}

Frontend: Simple Form

'use client';
import { useState } from 'react';

export default function Generate() {
  const [prompt, setPrompt] = useState('');
  const [response, setResponse] = useState('');
  const [loading, setLoading] = useState(false);

  const handleSubmit = async (e: React.FormEvent) => {
    e.preventDefault();
    setLoading(true);

    try {
      const res = await fetch('/api/generate', {
        method: 'POST',
        body: JSON.stringify({ prompt, model: 'mistral' }),
      });

      const data = await res.json();
      setResponse(data.response);
    } catch (error) {
      console.error(error);
      setResponse('Error generating response');
    } finally {
      setLoading(false);
    }
  };

  return (
    <form onSubmit={handleSubmit} className="space-y-4">
      <textarea
        value={prompt}
        onChange={(e) => setPrompt(e.target.value)}
        placeholder="Ask Ollama anything..."
        className="w-full border rounded p-2"
      />
      <button
        type="submit"
        disabled={loading}
        className="px-4 py-2 bg-blue-600 text-white rounded hover:bg-blue-700 disabled:opacity-50"
      >
        {loading ? 'Generating...' : 'Generate'}
      </button>
      {response && (
        <div className="p-4 bg-gray-100 rounded whitespace-pre-wrap">
          {response}
        </div>
      )}
    </form>
  );
}

Common Issues & Fixes

"Connection refused: localhost:11434"

Ollama not running. Start it:

./ollama serve

"FATAL: password authentication failed for user"

Wrong DATABASE_URL. Check Neon console | Connection string. Copy the full URL including ?sslmode=require.

"Error: Ollama connection timeout"

Model too slow or stuck. Check Ollama logs:

# If using ollama serve in background, check process
ps aux | grep ollama

Restart if needed:

killall ollama
./ollama serve

Memory issues on large models

Ollama supports quantized models (4-bit, 5-bit) to reduce size:

ollama pull mistral:7b-q4_0  # 4-bit quantized, ~4GB RAM

Deployment Options

Vercel + Neon + Local Ollama

Frontend + API routes | Vercel Database | Neon Ollama | Your machine (webhook triggers it)

This works but means requests timeout if Ollama isn't running. Better for batch jobs.

// app/api/queue/route.ts |--  Queue job to run later
export async function POST(request: Request) {
  const job = await request.json();
  // Store in Neon, have your local Ollama machine poll for jobs
  await pool.query('INSERT INTO jobs (prompt, status) VALUES ($1, $2)', [job.prompt, 'queued']);
  return Response.json({ id: job.id, status: 'queued' });
}

Docker + Compose

Run all three locally:

version: '3.9'
services:
  ollama:
    image: ollama/ollama:latest
    ports:
      - '11434:11434'
    volumes:
      - ./ollama-data:/root/.ollama
    environment:
      - OLLAMA_KEEP_ALIVE=5m

  web:
    build: .
    ports:
      - '3000:3000'
    environment:
      - DATABASE_URL=postgres://user:pass@postgres:5432/db
    depends_on:
      - postgres

  postgres:
    image: postgres:16-alpine
    environment:
      POSTGRES_PASSWORD: secret
    volumes:
      - ./postgres-data:/var/lib/postgresql/data

Example: vhsbox

vhsbox.netlify.app uses this exact stack to tag video metadata with Ollama, store in Neon, serve from Next.js. See the repo for full implementation.

Performance & Costs

| Scenario | Monthly Cost | Notes | | ---------- | ------------- | ------- | | 1000 requests/mo | $0 Ollama + $0 Neon + $0 Vercel | Free tier covers it all | | 10k requests/mo | $0 Ollama + $0 Neon + $0 Vercel | Still free; Neon free tier is generous | | 100k requests/mo | $0 Ollama + $15 Neon | Neon scales, Ollama is local | | 1M requests/mo | $0 Ollama + $300 Neon | Storage/compute adds up; consider splitting to OpenAI |

When to switch to OpenAI: If you hit 100k+ requests/mo and Neon costs spike, compare against OpenAI API. At that scale, trade-offs flip |-- API costs predictable, your Ollama machine becomes bottleneck.

Takeaway

Ollama + Neon + Next.js is the fastest path to a working AI-powered fullstack app with zero billing. Ship the MVP, measure demand, scale when needed. The pattern survives that transition |-- swap Ollama for OpenAI API, Neon for self-managed Postgres, all with one line changes.

Key Takeaways

  • Ollama runs LLMs locally with zero API costs and full data privacy
  • Neon provides serverless PostgreSQL with generous free tier
  • Next.js API routes eliminate need for separate backend service
  • Local-first development means faster iteration and offline work
  • Cost stays at $0 until you need production scale

Further Reading

Ollama Model Pull

ollama pull mistral

Thanks for reading!

Read More Articles