Building a Mobile-First AI Object Detector (Real-Time + Cloud AI)

by | Feb 13, 2026 | Articles, Projects | 0 comments

Introduction

In today’s world of computer vision and AI, building object detection applications has become more accessible than ever. In this article, I’ll walk you through how I built a modern, mobile-first AI Object Detector that combines:

  • ⚡ Real-time, browser-based detection
  • ☁️ Cloud-powered AI analysis

The Vision: Fast Detection When It Matters, Deep Analysis When You Need It

Most object detection apps fall into one of two categories:

  • Real-time detection, but limited detail
  • Deep AI analysis, but high latency

The solution? Both.

  • MobileNet (COCO-SSD) for instant, local detection
  • Vision Language Model (VLM) for optional deep analysis


Technology Stack

Frontend

  • Next.js 16 (App Router)
  • TypeScript
  • Tailwind CSS
  • shadcn/ui
  • Lucide Icons

AI / ML

  • TensorFlow.js
  • MobileNet COCO-SSD (80 object classes)
  • Z.ai VLM (Cloud AI analysis)

Key Features

  • 10–15 FPS real-time detection
  • Mobile-first responsive UI
  • Zero-latency local inference
  • Optional deep AI analysis
  • Offline-capable core
  • Graceful cloud fallback

Architecture Overview

flowchart TB
    UI["User Interface<br/>Camera · Stats · History"]

    UI --> LocalAI
    UI --> CloudAI

    LocalAI["Local AI (Browser)<br/>TensorFlow.js<br/>MobileNet COCO-SSD"]
    CloudAI["Cloud AI<br/>Z.ai Vision Language Model"]

    LocalAI --> Results
    CloudAI --> Results

    Results["Detection Results<br/>Objects · Confidence · Bounding Boxes"]

Implementation Details

1. TensorFlow.js Backend Initialization

import * as cocoSsd from '@tensorflow-models/coco-ssd'
import * as tf from '@tensorflow/tfjs'

const initializeTensorFlow = async () => {
  try {
    await tf.setBackend('webgl')
    console.log('✅ WebGL backend set')
  } catch {
    await tf.setBackend('cpu')
    console.log('✅ CPU backend set')
  }

  await tf.ready()
  return cocoSsd.load()
}

Why this matters

  • TensorFlow.js fails silently if no backend is ready
  • Explicit backend selection prevents runtime errors

2. Camera Handling (Video Ready State)

const startCamera = async () => {
  const stream = await navigator.mediaDevices.getUserMedia({
    video: {
      facingMode: 'environment',
      width: { ideal: 640 },
      height: { ideal: 480 }
    }
  })

  video.srcObject = stream

  await new Promise<void>((resolve) => {
    video.addEventListener('loadeddata', () => resolve(), { once: true })
  })

  await video.play()

  if (!video.videoWidth) {
    throw new Error('Video has no dimensions')
  }

  startDetectionLoop()
}

3. Real-Time Detection Loop

const startDetectionLoop = () => {
  let lastTime = Date.now()
  let frameCount = 0

  setInterval(async () => {
    if (!video || !model || video.readyState < 2) return

    const predictions = await model.detect(video, 10, 0.5)

    frameCount++
    const now = Date.now()

    if (now - lastTime >= 1000) {
      const fps = Math.round((frameCount * 1000) / (now - lastTime))
      setFPS(fps)
      frameCount = 0
      lastTime = now
    }

    setDetections(predictions)
    drawDetections(predictions)
  }, 100)
}

4. Bounding Box Rendering

const drawDetections = (predictions) => {
  ctx.clearRect(0, 0, canvas.width, canvas.height)

  predictions.forEach(p => {
    const [x, y, w, h] = p.bbox
    const score = Math.round(p.score * 100)

    const color =
      p.score > 0.8 ? '#22c55e' :
      p.score > 0.6 ? '#eab308' : '#ef4444'

    ctx.strokeStyle = color
    ctx.lineWidth = 3
    ctx.strokeRect(x, y, w, h)

    ctx.fillStyle = color
    ctx.font = 'bold 14px sans-serif'
    ctx.fillText(`${p.class} ${score}%`, x + 6, y - 7)
  })
}

5. Cloud AI Fallback Strategy

const analyzeWithAI = async () => {
  try {
    const res = await fetch('/api/detect-objects', {
      method: 'POST',
      body: JSON.stringify({ image })
    })

    if (!res.ok) throw new Error()
    setResult(await res.json())

  } catch {
    setResult({
      source: 'local',
      objects: detections.map(d => d.class),
      analysis: detections
        .map(d => `- ${d.class} (${Math.round(d.score * 100)}%)`)
        .join('\n')
    })
  }
}

Mobile-First Design

Responsive Layout

<div className="grid lg:grid-cols-3 gap-6">
  <div className="lg:col-span-2">Camera View</div>
  <div>Stats Panel</div>
</div>

Touch-Friendly Controls

<Button
  size="lg"
  className="min-h-[48px] px-8"
>
  Capture
</Button>

Performance Optimizations

  • Confidence threshold: 0.5
  • Detection interval: 100ms
  • Model loaded once and reused
  • WebGL preferred on supported devices

Supported Object Classes (COCO-SSD)

  • People
  • Animals
  • Vehicles
  • Furniture
  • Kitchen Items
  • Sports Equipment
  • Outdoor Objects

(80 total categories)


Future Enhancements

  • 🌍 Multi-language detection
  • 🎯 Custom-trained models
  • 🔊 Audio feedback
  • 📈 Object tracking
  • ☁️ Cloud history storage

Conclusion

This architecture delivers:

  • ⚡ Speed when it matters
  • 🧠 Intelligence when needed
  • 📱 Excellent mobile performance
  • 🛡️ Graceful failure handling

Client-side AI + Cloud AI = Best of both worlds


Try It Yourself

The full source code is available on GitHub.
Fork it, customize it, and build your own AI-powered vision apps 🚀

Written by

Related Posts

0 Comments

Submit a Comment

Your email address will not be published. Required fields are marked *