AI News

  • Loading...

How to Use Ask Photos AI on Google Photos: A Complete Guide

How to Use Ask Photos AI on Google Photos: A Complete Guide

Ask Photos is a new AI-powered search feature in Google Photos, backed by Google's Gemini technology. It lets you find images and information using natural language queries instead of traditional keywords. What's interesting here is that Google has embedded Gemini directly into Photos, fundamentally changing how you interact with your library.

Google Photos already had solid search capabilities, but Ask Photos takes things to another level. The AI-driven approach gives you more options and better results overall. Instead of typing keywords, you can now describe what you're looking for in plain English—and Gemini understands the context behind your query to surface exactly what you need. Let's walk through how to get started with Ask Photos.

How Does Ask Photos Work?

Ask Photos is Google Photos' new AI search tool that understands descriptive queries and returns relevant results. Under the hood, it uses Google's Gemini AI models to analyze context and deliver smarter search results.

The feature understands a ton of information stored in your Google Photos library—important people in your life, favorite meals, locations you've visited, hobbies, and much more. You can query it with natural questions or descriptions, adding as much detail as you want.

Because it grasps context, Ask Photos can surface images that traditional keyword searching would struggle to find. The real power comes from its ability to connect the dots across your entire photo collection.

How to Use Ask Photos AI on Google Photos

Step 1:

Open Google Photos and you'll see a notification about the new AI search tool. Accept it to enable the feature. Next, tap the search icon to activate Ask Photos. You'll now see a search bar where you can enter your description.

Step 2:

Type in a description of the image you're trying to find. Ask Photos understands natural language, so it will process your query and return related results. The more descriptive you are, the better your results.

You can continue entering new search queries to find different images. Gemini scans your photos and identifies matches based on the details in your description. The results align closely with your search terms, so specificity pays off.

The results Ask Photos delivers will closely match your search criteria. That's why detailed descriptions yield better outcomes.

Ask Photos Google Photos

Step 3:

Gemini has its own settings panel within Google Photos so you can tweak how it operates. In the Ask Photos search interface, tap the three-dot menu in the top right corner and select Gemini in Photos.

Adjust Gemini in Google Photos

You'll now see Gemini settings for Google Photos. Review these options and disable any features you don't want to use.

Gemini settings in Google Photos

What Can You Do With Ask Photos?

Ask Photos, powered by Gemini in Google Photos, lets you search for images using anything from simple to highly specific questions. The AI grasps context to locate the right photos. Here are some practical search examples:

  • Find travel photos: Ask "Where did I go to [location]?" or "Show me photos from my trip to [location] in [year]." Ask Photos pulls from multiple data points to find the right images.
  • Find food photos: Forgotten what you ate? Try "What did I eat at the resort in [location]?" or "Show me all the desserts I had this year."
  • Find specific details: You can ask hard-to-remember questions like "What's my license plate number?" Ask Photos searches through related images to provide the answer.
  • Find particular photos: Short queries like "Recent selfies" or "Landscape photos" work too, helping you quickly locate specific image types.
  • Find photos with specific people: Ask "When was the last time I saw my sister?" or "Photos of me and my brother at the restaurant." Thanks to its face recognition capabilities, Google Photos pinpoints relevant images with impressive accuracy.

Description: Learn how to search your photo library using natural language with Ask Photos, Google's new AI-powered search feature powered by Gemini.

Related Articles

Three Major Shifts That Will Define AI's Future Over the Next Year

Three Major Shifts That Will Define AI's Future Over the Next Year

Everyone loves saying AI is moving at lightning speed, but most AI commentary today confuses noise with real signal. The three shifts below are far more foundational: they're unfolding relatively quietly, but once they take hold, they'll reshape everything built on top. Over $1 trillion has been raised and spent to get AI where it is today. In the next 12 months, three core transitions will determine where that massive sum actually creates value.

Shift #1: The Real Value Lives in the Harness, Not the Model

For the past two years, the AI world has been obsessed with one question: whose AI model is the smartest? That's actually the wrong question. What we really need isn't a smarter model—it's a better system.

Prediction 1: Harness Will Become the Key Asset

Developers are starting to build loyalty around harness rather than specific AI models, the way a programmer might stick with Vim or Emacs.

Large language models process text. Fundamentally, they're probabilistic prediction systems without any guarantees. A harness is deterministic code wrapped around your interaction with the model. It contains business logic, evaluation loops, context management, and security guardrails.

Advanced harnesses add safety features, use-case-specific logic, and even complex workflows. Developers are adopting specialized harnesses for coding—tools like Claude Code or OpenCode—that let them swap models based on preference or cost while keeping a high-quality workflow intact. Task-specific harnesses are also emerging for presentations, code generation, document editing, and data collection.

Think of an LLM like an untrained employee: they don't need to be brilliant if you give them a well-designed process to follow.

Prediction 2: Companies Will Build Their Own Harnesses

We can probably stop treating prompt engineering as a long-term skill. Enterprise AI development is essentially reverting to what software engineering has always been: converting business processes into deterministic code.

Non-deterministic models let you scale heuristics and handle complex edge cases—which are actually the real source of technical debt and complexity in any system. An LLM that correctly handles 90% of edge cases is already a fantastic tool. To leverage that, companies will build custom harnesses designed around their proprietary business logic, using frameworks like LangGraph, CrewAI, or AutoGen.

Generic harnesses from Anthropic and Google work fine for routine knowledge work. But as AI enters a new "industrial revolution" at scale, companies will need enterprise harnesses built in-house. This is the inevitable direction the industry is heading.

Prediction 3: Document-Based Workflows Will Become the Standard

Single-agent systems are easy to deploy, but when you want to scale both power and complexity, you move to multi-agent workflows. Just like in the real world, companies aren't collections of totally independent individuals. They're teams, departments, and organizations where each person handles specific tasks but coordinates to complete multi-person processes.

Like humans, agents work most efficiently when focused on one task at a time. Each agent can be a specialist, and the system scales by connecting them together. There are plenty of complex frameworks for this—just like there are countless task management systems for humans. But companies keep gravitating toward document-based workflows.

From wikis to network drives, most companies already run on documents, folders, and forms. Each form is like an API for humans—it feeds data into a process. Each updated document acts as an event that moves the process to the next step.

Document-based workflows are popular for a simple reason: they're easy to deploy, easy to modify, and genuinely scalable. They provide visibility and auditability, support security and access control, and integrate naturally across teams—even across organizations.

Sure, agent-to-agent orchestration systems and powerful workflow managers will keep evolving. But simple document-based workflows will become the default tool. They let agents coordinate with humans and are the fastest path to dropping AI into existing processes—because most companies already have the infrastructure they need.

This is exactly what agents were built for. They're particularly good at text, documents, reading, and filling forms. The tedious parts of document work that humans hate doing? That's where agents excel. Injecting agents into routine, repetitive knowledge work is where we'll see real transformation happen.

Shift #2: AI Models Are Becoming Interchangeable Parts

As value shifts from model to harness, the underlying foundation model becomes just another replaceable component—like a part in a machine. And when something becomes interchangeable, its price drops.

Prediction 4: The Race for Cutting-Edge LLMs Will Cool Off

The last two years have been driven by a frontier model arms race, with major providers constantly releasing more powerful models (except where the US blocks them).

That race is now hitting economic reality. Billions in subsidies are giving way to actual API bills. We're already seeing pricing move closer to real costs—along with plenty of stories about shocked CEOs facing massive AI invoices or companies rehiring humans to manage expenses.

What's interesting here is that despite models getting stronger, actual practical value isn't increasing proportionally. Smaller, cheaper models are good enough for most tasks. Just like everything else—not every car needs to be luxury, not every CPU needs to be the fastest. Most of the world runs on cheap components, and most of AI's future will too.

Prediction 5: Open-Source LLMs Will Win

Cheap components you don't control won't stay cheap for long. That's the other half of the commoditization story.

Once vendor lock-in takes hold, cheap components you can't control quickly become expensive. Companies building custom harnesses have a massive structural advantage: model independence.

Cloud-hosted models are currently heavily subsidized, making token costs far lower than actual expenses. Providers are betting that LLM operating costs will drop fast enough to keep the market viable while giving companies time to migrate workloads to hosted solutions. But once you're locked into a platform, prices will rise. Many tasks you'd happily hand to AI today might not pencil out on those platforms anymore.

Technology only truly scales when it becomes a commodity. That's why serious AI companies will lean on open-source models for most daily needs. Routine tasks don't need frontier models. Major companies like Netflix are already running their own LLM servers with open models to control costs more directly.

That's where custom internal harnesses become critical. Companies use them to build deterministic guardrails—and one of the most important guardrails is cost control.

Shift #3: LLMs Are Just One Piece of a Larger, Connected AI System

If models are just a commodity component, the real system will be a combination of many different technologies, where multiple types of AI work together and increasingly operate across organizational boundaries.

Prediction 6: LLMs Won't Be the Silver Bullet

When relational databases arrived, people wanted to use them for everything: message queues, file systems, you name it.

Generative AI isn't a silver bullet either. I've seen teams use LLMs for recommendation systems, predictive analytics, machine learning pipelines, knowledge systems, search tools, and data analysis. Like a general-purpose database, LLMs adapt remarkably well. But they're also expensive and inefficient for some jobs.

Generative AI is built on decades of existing AI techniques—deep learning, machine learning, natural language processing, predictive models. Those foundational technologies still exist, and for many tasks they outperform a general-purpose LLM.

When companies deploy custom harnesses, they'll build multi-modal agents using the entire spectrum of AI technology. LLMs will handle the fast, flexible coordination between complex processes, but they'll work alongside other AI systems instead of replacing them. LLMs won't solve every problem.

Prediction 7: Today's Protocols Are Just the Beginning

MCP is a dead protocol for the internet. I'm not sure whether A2A will be the replacement, but agents will definitely operate across organizational boundaries and need safe coordination.

Internet history is a story of gradually retreating from an overly open system, then layering on security, authentication, and observability later. MCP is following that same path.

Given current opportunity scale and hard-won lessons, agent-to-agent coordination at internet scale will start with a protocol designed from the ground up for security and scalability, rather than bolting these on afterward.

What Does This Future Look Like?

Stitch all three shifts together and the picture gets clear. Value will flow away from models and into harnesses. Models, no longer the prize, will become a commodity that companies should run themselves instead of renting from others. And what we build with those models won't be a single all-powerful AI—it'll be a connected, multi-component system of specialized AI experts orchestrated through one of the simplest and most resilient infrastructures we have: documents.

Every technology platform eventually commoditizes at the foundation and creates differentiation at the top. CPUs became a commodity. Linux became a commodity. Cloud computing did too. Business value then moves up to applications, workflows, and data. LLMs are following the same path.

This may not match the vision AI hype promised over the past few years. But it might actually be better, because this is AI actually getting deployed and creating real value. Teams that understand these shifts, invest in harnesses, use cost-manageable models, and deploy agents to automate routine, repetitive business work will pull ahead.

Teams still waiting for the next generation of AI models to "save" them? They'll probably be waiting a while.


Description: The AI industry is confusing noise with signal. Here are three fundamental changes that will actually matter.

Related Articles

Hybrid AI Architecture: Combining RAG and Fine-Tuning for Enterprise-Grade Support Systems

Hybrid AI Architecture: Combining RAG and Fine-Tuning for Enterprise-Grade Support Systems

Companies today face a tough challenge: they need chatbots that are secure, accurate, and fast. More specifically, they need AI systems that handle support questions reliably without exposing sensitive data or sounding like a robot. The stakes are real. Data breaches, slow response times, and poor answer quality directly impact the bottom line. According to IBM's 2025 report, the average cost of a data breach globally reaches $4.44 million. A single poor customer interaction can erode trust quickly—and that trust is hard to rebuild.

But here's the problem: off-the-shelf chatbots and large language models (LLMs) rarely meet enterprise expectations. LLMs are powerful, sure, but they come with real limitations—token limits, difficulty leveraging context effectively, and hallucinations (making up confident-sounding answers). These weaknesses become obvious when you need domain expertise, strict response formats, and guardrails. So the question becomes: How do you build an AI system that reasons like an expert, grounds answers in actual data, runs fast, stays secure, and remains controllable?

After years of deploying AI models, one thing is clear: no single technique solves this. What you really need is an architectural approach that separates what the model knows from how it generates answers, while combining learning capabilities with retrieval mechanisms.

The Four Core Challenges

Let's dig into the real constraints you face.

Context Windows Aren't as Big as They Sound

Modern LLMs advertise 16,000, 32,000, or even 128,000 token context windows. In practice? Anyone who regularly uses these models knows attention quality degrades much faster. Load a large chunk of text into the context and the model struggles to effectively use information in the middle of the prompt—a phenomenon called primacy-recency bias. It overfocuses on what comes first and last, missing the critical stuff in between.

Simply expanding the context window doesn't guarantee better answers. In enterprise settings where your knowledge base contains millions of tokens, this isn't a viable solution.

Models Miss Uncommon Information

Even when relevant information sits right there in the prompt, LLMs overlook it, misinterpret it, or weigh irrelevant parts too heavily. The "Lost in the Middle" research backs this up—long context inputs lead to incomplete reasoning unless carefully controlled.

This breaks the "throw everything into the prompt" strategy, especially for customer support systems managing large, complex domain knowledge.

Retrieval Creates a Latency-Accuracy Trade-Off

Pulling information has computational costs. Retrieve too much data and response time crawls while the model's focus deteriorates. Retrieve too little and hallucination risk climbs. The real challenge isn't whether to retrieve—it's retrieving the right amount: enough context to ensure accuracy without overloading the system or the model.

Hallucinations When Context Runs Out

LLMs rarely refuse to answer. Instead, they confidently invent information that sounds plausible but is entirely made up. For customer support, this is unacceptable—it destroys credibility, introduces errors, and creates compliance nightmares.

These four problems point to the same conclusion: adding more context isn't the answer. You need smarter architecture.

The Solution: Hybrid AI Architecture

Combining Retrieval-Augmented Generation (RAG) with fine-tuned language models works. The key insight? Fine-tuning and retrieval solve different problems.

Fine-tuning teaches the model how to answer. Retrieval feeds the model information to answer with. Forcing one technique to do both jobs leads to inefficiency, instability, or high costs. The smarter approach: let each component play to its strengths.

Using RAG to Boost Accuracy Through Smart Retrieval

Don't dump raw documents into the model. Instead, build a searchable knowledge base curated from internal Q&A pairs, product guides, technical docs, and policy references. When the system processes a query, the retriever selects only the most relevant chunks and feeds them into the prompt. Answers stay grounded in verified data.

This cuts hallucinations significantly, improves accuracy, and speeds up responses by keeping the context window tight and focused. But RAG alone falls short. Even with perfect retrieval accuracy, outputs vary wildly in tone, structure, format, and instructional detail.

In a real case study, a specialized customer support chatbot had access to nearly 100% of the correct context—yet answer accuracy barely hit 70%. The small language model struggled to extract meaning from long contexts and couldn't maintain the conversational tone needed to guide users toward deeper technical discussions or follow-up meetings.

This reveals RAG's core limitation: it provides information, but can't teach the model how to reason or communicate in a specific domain.

Fine-Tuning Qwen: Teaching the Model Communication Style

To improve consistency, tone, and reasoning, the team fine-tuned Qwen on roughly 1,000 carefully selected expert Q&A pairs in their target domain. The goal wasn't to add new facts. Instead, the model needed to learn specialized language, maintain brand voice, follow consistent response formats, reason step-by-step, and handle edge cases in the support process.

Fine-tuning adjusts how the model behaves, not what it knows. This distinction matters. Full fine-tuning risks catastrophic forgetting and demands enormous compute. To avoid this, they used Low-Rank Adaptation (LoRA) adapters, which adjust only small adapter matrices while preserving the base model's general knowledge. It cuts GPU memory requirements while maintaining near-equivalent performance.

Results improved noticeably. The model became more consistent and nuanced. For stable, routine questions, it delivered accurate answers repeatedly without extra retrieval. But as expected, it struggled with new features, updated policies, or rare information.

Back to the chatbot example: fine-tuning pushed tone accuracy to around 90%, but information accuracy dropped to 50%. The lesson reinforced: fine-tuning cannot replace retrieval.

Why RAG or Fine-Tuning Alone Falls Short

The experiments expose the trade-off clearly. RAG-only systems excel at staying faithful to real data and updating with fresh information, but suffer from inconsistent tone and higher latency. Fine-tuning-only systems shine in tone and response structure, but struggle when knowledge changes or rare facts appear.

Picking just one means accepting the other's weaknesses. Combine a fine-tuned model with RAG and you beat both approaches independently. Tone accuracy reaches about 75%—better than RAG alone (which has no tone control) yet lower than fine-tuning-only (90%). Meanwhile, information accuracy hits around 73%—better than fine-tuning-only (50%) and raw RAG (70%).

What's interesting here is that domain understanding and output formatting learned during fine-tuning help the model apply retrieved information more effectively than an untuned base model could. In other words, RAG provides accurate, current knowledge while fine-tuning teaches the model how to use that knowledge to craft appropriate responses. That's why hybrid architecture balances accuracy, consistency, and performance better than either technique alone—exactly what enterprise support systems need.


Description: Learn how RAG and fine-tuning work together to build accurate, reliable AI assistants that balance speed, consistency, and real-world knowledge.

Related Articles

Beyond Code: How to Use AI Coding Agents for Non-Technical Tasks

Beyond Code: How to Use AI Coding Agents for Non-Technical Tasks

AI Coding Agents were designed to help developers—fixing bugs, adding features, writing new code. But if you only use them for that, you're leaving serious potential on the table. The truth is, tools like Claude Code, Cursor, Codex, and Gemini CLI are remarkably capable at handling all sorts of non-coding work on your computer. And once you start thinking differently about what they can do, you'll wonder how you ever managed without them.

You could hand off auto-filling web forms, managing your personal budget, finding the cheapest product prices, researching customer data, preparing reports, or even learning a new language. But the real shift is mental: instead of assuming you have to do everything yourself, start asking "Could an AI Coding Agent handle this for me?" constantly.

Why Use Coding Agents for Everything You Do on Your Computer?

Most people assume Coding Agents are exclusive to programmers. That misconception misses the bigger picture. These tools can handle nearly any task you perform on a computer. They don't just understand code—they can read documents, interact with web browsers, manipulate files, search for information, and automate countless workflows.

When we talk about Coding Agents here, we mean Claude Code, Cursor, Codex, Gemini CLI, and similar platforms. Each has distinct strengths, but overall they're becoming smarter every update and can take on work that previously required manual human effort.

There are two immediate advantages. First: time savings. By running multiple tasks in parallel through AI, you accomplish more in the same hours. Second: automation of repetitive or tedious work, freeing you to focus on what actually matters.

Non-Technical Work That Coding Agents Can Handle

Here are task categories you can genuinely delegate to Coding Agents regularly:

  • Personal budget management and expense tracking.
  • Creating HTML reports.
  • Designing presentation slides.
  • Accessing and interacting with websites.
  • Supporting sales operations.
  • Researching information online.
  • Learning foreign languages.

But the list isn't what matters—the mindset is. Every time a new task lands, ask yourself whether a Coding Agent could help.

Budget Management, Reports, and Presentations

One of the most practical applications is personal finance management. Download your transaction history from your bank or digital wallet, feed it to a Coding Agent, and let it work. The AI categorizes income and expenses, creates a budget, and generates an HTML financial report you can actually visualize and track over time.

Creating presentation slides gets dramatically easier too. Describe your topic, and a Coding Agent generates a complete PDF presentation immediately. Then you just request refinements until it matches your vision.

Here's an effective pro tip: give the Coding Agent samples of presentations or documents you've created before. It learns your style, color choices, layouts, and writing voice—then applies that "personal stamp" to everything new it creates. This trick works for countless other tasks. When AI sees your old work, it understands your patterns far better.

Let Coding Agents Handle Website Interactions

Another genuinely useful application: delegating website tasks to a Coding Agent.

You can set up a dedicated browser profile for your Coding Agent with your necessary accounts pre-saved. Then the AI logs in automatically, creates API keys, changes settings, downloads data, or performs other actions—all without you lifting a finger. While it works in the background, you tackle other priorities. This is especially valuable for complex websites with many steps or confusing interfaces.

One caveat: some websites explicitly prohibit AI agents in their terms of service, or deploy CAPTCHAs and verification measures specifically to block automated access.

Sales Support and Research

Coding Agents make effective sales assistants. If you work with customers or manage a CRM, create a dedicated folder containing all your sales data so the AI always has full context when working on your behalf.

With everything stored centrally, the Coding Agent remembers interaction history, understands your customers more deeply, and becomes increasingly effective. It can independently research prospect information, track progress, set reminders, and consolidate data from multiple sources.

Beyond sales, AI excels at online research. Say you want blackout curtains in a specific size at the lowest price. Give your requirements to a Coding Agent. It searches the internet, compares options across retailers, builds a price comparison table with links, and delivers everything in minutes. Instead of opening dozens of websites yourself, you get a complete summary ready to review.

Using Coding Agents as Language Tutors

Outside of work, Coding Agents become personal language tutors. Create a dedicated learning folder where the AI maintains your vocabulary list, notes which words you tend to forget, schedules 10-minute daily review sessions, and tracks your progress over time.

What's interesting here is that modern language models support dozens of world languages. They converse fluently in uncommon languages like Norwegian. While slang and regional dialects sometimes trip them up, their teaching and conversation abilities remain genuinely strong.

Shift Your Mindset Around Coding Agents

The most important change isn't learning new features—it's rethinking how you approach work.

View Coding Agents as a collaborator present for every computer task. Rather than working solo, AI backs you up by recalling forgotten information, remembering workflows so it handles them independently next time, and taking repetitive work off your plate. You accomplish more simultaneously and productivity rises noticeably.

Obviously, not everything should go to AI. Personal communication—responding to close friends or family—deserves your human touch. Don't use AI to write personal essays or intimate social posts. Your writing reflects your thinking and perspective. Handing that entirely to AI drains the meaning from what you've created.

Coding Agents Enhance Rather Than Replace

Some argue that creative work shouldn't be delegated to AI. That view oversimplifies things. Coding Agents are tools that expand human capability. You still set the goals, provide the requirements, and direct the AI's work. When it runs automatically, it's executing what you designed.

Actually, Coding Agents can boost creativity. Maybe you've had dozens of app ideas but lacked time or skills to build them. Now a Coding Agent quickly transforms those concepts into working prototypes. Ideas that stayed on paper before now become real projects worth testing.


Coding Agents aren't exclusively for developers. Any task performed on a computer has potential to be supported or automated by AI at some level.

The real habit to build: constantly ask "Can I delegate this to a Coding Agent?" Whether it's financial planning, report writing, slide design, market research, or language learning, today's AI Coding Agents save considerable time and effort.

You remain the final decision-maker. Coding Agents don't replace you—they serve as an intelligent assistant, helping you work faster, accomplish more, and reclaim hours to spend on what genuinely matters.


Description: AI Coding Agents aren't just for developers. Learn how tools like Claude, Cursor, and Gemini CLI can automate everyday computer work.

Related Articles

Rowboat: The Open-Source AI Assistant That Does What Claude Cowork Should Have Done From Day One

Rowboat: The Open-Source AI Assistant That Does What Claude Cowork Should Have Done From Day One

Using an AI assistant has always felt fragmented. Sure, AI can help you write that email, but you're still the one shuttling information back and forth—copying text into ChatGPT or Claude, explaining what you need, waiting for a draft, then pasting it back into Gmail. It's tedious. What's frustrating is that the AI isn't really the bottleneck; it's all this manual context-switching between tools. The good news? That friction is starting to disappear, thanks to products like Claude Cowork. But there's an open-source alternative that delivers something many people actually wish Cowork had built in from the start.

Download Rowboat for Windows Download Rowboat for macOS Download Rowboat for Linux

Rowboat Is the Open-Source Answer, With Persistent Memory Built In

Rowboat's main dashboard showing an evening greeting, a gym event scheduled for 11 hours later, and 3 new messages in the inbox
Rowboat's main dashboard showing an evening greeting, a gym event scheduled for 11 hours later, and 3 new messages in the inbox

Let's back up for a moment. Claude Cowork essentially takes the agentic capabilities that power Claude Code and repackages them into something friendly for non-technical users. Instead of living in a terminal window focused purely on coding, Cowork gives Claude a proper workspace where it can manipulate files, integrate with other services, and handle multi-step workflows across different use cases.

Rowboat pursues a similar "AI coworker" concept but from a different angle entirely. The best way to think about Rowboat is as an AI workspace that continuously builds up knowledge about your work in the background. Rather than forcing you to rebuild context every time you need help with something new, Rowboat maintains a persistent knowledge graph about the people, projects, conversations, meetings, and information that make up your professional life.

Rowboat is completely open-source, which means you can see exactly how it works and modify it if you want. It also isn't locked to any specific AI model, so you can run local models through Ollama or LM Studio, or plug in any model hosted on your preferred service. Since Rowboat is open-source and runs on your own machine, your data stays local—stored as regular Markdown files instead of getting locked into some proprietary cloud service.

Rowboat Feels Like a Genuine Workspace, Not Just a Chatbot

Rowboat's main interface showing an inbox preview and a dashboard of background AI agents
Rowboat's main interface showing an inbox preview and a dashboard of background AI agents

Rowboat feels intentionally designed for work that doesn't fit neatly into a single session. Everything connects to the same system. Your email, calendar, notes, meetings, browser activity, and conversations all feed into the system's "Brain." That accumulated context stays ready whenever you return to work later.

You'll notice this immediately when using Rowboat to plan an upcoming business trip. It doesn't just dig through forwarded emails and extract dates—it cross-references them with your Google Calendar, identifies what's confirmed versus tentative, spots gaps in your plan, and reminds you of pending meeting invitations. Then when you ask follow-up questions about specific events in a separate chat thread, you don't have to explain the entire trip context all over again.

A Rowboat conversation titled 'Germany Business Trip Itinerary' showing notes about Turkish Airlines bookings and calendar entries
A Rowboat conversation titled "Germany Business Trip Itinerary" showing notes about Turkish Airlines bookings and calendar entries

What's interesting is that Rowboat genuinely feels like a workspace rather than a tool you visit when you need something done. The main dashboard works as a central hub for your work: upcoming calendar events, new emails, unfinished tasks, background processes, plus access to your notes, meetings, Brain, and separate workspaces—all without leaving the app.

The interface actually resembles those Notion dashboards people spend hours manually setting up. The crucial difference? Rowboat doesn't just show links to information. The inbox, calendar, notes, and projects displayed on the interface actively feed contextual data back to the AI. Everything connects by design from the start, rather than being bolted on as separate integrations to a chatbot.

You Can Use Almost Any AI Model With Rowboat

One of Rowboat's biggest advantages over Cowork is flexibility. With Cowork, you're locked into Claude by default. Rowboat lets you choose which AI model powers your experience. Use Rowboat's built-in models (with higher quotas on paid plans), or configure your own. It supports OpenAI, Anthropic, Gemini, OpenRouter, local models via Ollama, and any custom endpoint compatible with OpenAI's API.

Rowboat's "Brain" Feature Is More Impressive Than It First Appears

Rowboat Brain's graph view showing 20 interconnected notes across different folders
Rowboat Brain's graph view showing 20 interconnected notes across different folders

If you've ever used Obsidian's graph view, Rowboat Brain will feel instantly familiar. As Rowboat learns about your work, it creates linked notes about people, companies, projects, and topics it encounters, gradually building a visual knowledge graph. Click any node and you see the actual information Rowboat has collected and stored about that entity.

At first, many people assume it's just a pretty visualization of information extracted from email. Then connect your Gmail account and let Rowboat run. Within surprisingly little time, it builds a profile about you with details that go far beyond basic facts. For instance, if you're a tech journalist, Rowboat recognizes that immediately—it knows which publications you write for, what kinds of products you review, and even some of the PR relationships you've built. That's not shocking given it has email access.

A Rowboat Brain profile page for Mahnoor Faisal, a tech journalist working in Karachi, Pakistan
A Rowboat Brain profile page for Mahnoor Faisal, a tech journalist working in Karachi, Pakistan

What's genuinely surprising is the subtlety and nuance in how it analyzes. Rowboat dedicates entire sections to describing how you communicate with external PR contacts. It notices you typically open with friendly greetings and thanks, keep responses concise, mention which publication you represent when relevant, and provide specific time slots when scheduling press calls or information exchanges.

Then it goes deeper still. Rowboat creates a separate section for long-term PR relationships and spots how your writing style shifts when talking to people you know well. It observes that you become more casual, throwing in "lol," "hahaha," or capital letters for emphasis, and that you're generally more open and enthusiastic when sharing feedback about products you've tested with familiar contacts.


Description: Meet Rowboat, an open-source AI workspace that maintains continuous context about your work. Here's how it compares to Claude Cowork.

Related Articles

Building and Managing Skills in Gemini Spark: A Complete Guide

Gemini Spark is a powerful feature within Gemini designed to automate workflows by defining specific processes and instructions. Here's what makes it useful: instead of repeatedly typing the same commands or entering identical parameters for similar tasks, you can build a Skill once—complete with guidelines, rules, and reference materials—and Gemini will apply it automatically whenever you need it. What's interesting here is that you can chain multiple Skills together to create sophisticated workflows that handle everything from simple, one-off tasks to complex, multi-step operations. Below, we'll walk you through creating and managing Skills in Gemini Spark.

How to Create Skills in Gemini Spark

Building a Skill through Gemini's AI Assistant

Step 1:

Start by opening Gemini and selecting Spark. Once you're there, scroll down and choose Skills from the menu.

Next to that section, you'll see Create with Gemini. Click on it to get started.

Step 2:

You'll transition to Gemini's interface, where the AI assistant helps you build your Skill. Describe the Skill you want with as much detail as possible so Gemini can properly implement it. Type your request in the input box below.

Instantly, your Skill springs to life based on your description. Below it, you'll see the command syntax for invoking this Skill.

Further down, you'll find a step-by-step workflow outline that breaks down exactly how Gemini will execute the Skill based on your description.

Using Pre-Built Templates to Create Skills

In the Skills section, under Suggestions, you'll discover ready-made templates you can use immediately.

Each template includes the core Skill content and step-by-step instructions. To create the Skill, click Create in the top right corner. You can also use these templates as examples to guide your own custom Skill development.

After a moment, Gemini Spark saves your new Skill, ready for whenever you need to use it.

Creating a Skill from Scratch with a Blank Template

In the Skills section, click on Create Manually to build a Skill entirely from scratch.

You'll see a form where you enter the Skill name at the top, followed by a section for detailed instructions on how the Skill should work. Once you've filled everything in, click Create to save it.

Managing Your Gemini Spark Skills

The Skills interface displays all the Skills you've created in Gemini Spark. To access management options for any Skill, click the three-dot menu icon next to it. A dropdown menu appears with various options.

To use a Skill right away, select Use Now. You can also edit the Skill—either with Gemini's assistance or manually.

When you select Use Now, the Skill appears in chat as /skillname. You then input your specific request, and Gemini executes the Skill based on the instructions and guidelines you defined earlier. Gemini returns the output matching your request—for example, below is a Skill designed to generate marketing ideas for bottled coffee products.

The complete output follows the instructions you created for the Skill.

Skill execution result in Gemini Spark


Related Articles

What is Meta AI on Threads? Complete Guide to Using It

Meta just rolled out a fresh feature on Threads called Ask Meta AI, and it's designed to boost adoption of Meta's chatbot while making life easier for users hunting down information about posts. Here's what's interesting: you can fire questions at the bot about literally any aspect of a post—whether that's summarizing the text, digging into embedded videos, or even requesting information Meta pulls from external sources. All without leaving the app.

Meta is rolling this AI-powered platform across its entire ecosystem in different forms. On Instagram, it powers image editing tools. On Threads, it works as an inline chatbot for individual posts. Users can accomplish a lot through simple, straightforward commands—no complex prompts required. Below is your step-by-step guide to getting the most out of Ask Meta AI on Threads.

How to Use Ask Meta AI on Threads

Step 1:

Open any post on Threads and tap the three-dot menu icon in the top right corner. Then select Ask Meta AI from the menu.

Step 2:

The AI chat interface will open and automatically display key information about the post—including text, images, videos, and other media. No input needed on your part. Meta serves up the essential details in your language, highlighting the most important points from the post.

Want deeper insights or have questions about video content? Just type your query and send it.

Step 3:

Meta AI extracts and analyzes video content for you. Everything comes back in your preferred language, even if the original post or video was created in a different language.

Need a condensed version? Simply ask Meta AI to create a summary of the key points.

What Can You Do With Ask Meta AI on Threads?

Ask Meta AI lets you get quick answers about any post directly within Threads. No need to copy-paste content into a separate chatbot. The real benefit here is convenience—everything happens in context.

The feature supports several practical uses:

  • Summarize lengthy posts when you don't have time to read the full text.
  • Break down confusing terminology or complex concepts mentioned in the post.
  • Get answers to follow-up questions related to the post's topic.
  • Receive additional context or background information for better comprehension.
  • Translate post content into another language.
  • Interpret images attached to posts, when AI image analysis is supported.


Related Articles

5 Essential Books to Master Large Language Models

5 Essential Books to Master Large Language Models

Generative AI is evolving at breakneck speed, but here's the thing: the mathematical principles and architectural foundations behind large language models (LLMs) have actually been well-documented for years. A few years ago, understanding traditional NLP and RNNs was enough. Then Transformers arrived and reset the game entirely. Now, grasping how these models are built, trained, and deployed has become essential knowledge for AI engineers, data scientists, and anyone shipping AI-powered products.

If you're ready to move beyond just calling ChatGPT or Claude APIs—if you actually want to understand how a foundation model works—you need structured learning material, not scattered blog posts. Here are five highly-regarded books that will take you from theoretical foundations all the way through practical implementation of Large Language Models.

1. Build a Large Language Model (From Scratch) – Sebastian Raschka

The best way to understand a complex system? Build it yourself. That's the philosophy driving Build a Large Language Model (From Scratch) by Sebastian Raschka.

Rather than just explaining concepts, this book walks you through constructing, training, and fine-tuning an LLM from the ground up using PyTorch. Every step is detailed thoroughly, giving you hands-on exposure to core components like tokenization, embeddings, attention mechanisms, Transformer architecture, and training optimization techniques.

What's interesting here is the book comes with over 20 annotated Jupyter notebooks. You're not just reading—you're coding alongside, watching data flow through each layer of the Transformer network. Perfect for AI engineers and researchers who need to truly understand what happens at each computational step. The theory-plus-code combination is particularly powerful.

2. The Hundred-Page Language Models Book – Andriy Burkov

Not everyone has time to build models from scratch. If you need a comprehensive overview fast, The Hundred-Page Language Models Book by Andriy Burkov is your move.

The title's honest: it compresses the entire evolution of language models into roughly 100 pages without sacrificing technical accuracy. Burkov guides you from classic n-gram models through modern architectures like BERT and GPT. Mathematical concepts are explained clearly with visual diagrams and concise Python examples throughout.

Each topic—pretraining, attention, text generation—gets its own focused chapter. You won't feel overwhelmed. This is ideal for students or working professionals who want solid foundations without a massive time commitment. The breadth here is impressive for the page count.

3. Hands-On Large Language Models – Jay Alammar & Maarten Grootendorst

Once theory is solid, the next phase is applying LLMs to real problems. Hands-On Large Language Models is written exactly for that transition.

Jay Alammar built his reputation on crystal-clear Transformer visualizations. Maarten Grootendorst is a seasoned NLP specialist. Together, they created something balanced: theory meets practice.

The real standout? Over 250 illustrations. Concepts like attention heads and multi-layer Transformer architecture become intuitive when drawn well. The book also covers semantic search, dense retrieval, prompt engineering, and RAG (Retrieval-Augmented Generation) systems. You'll learn fine-tuning and deployment using open-source tools, especially the Hugging Face ecosystem. This one's perfect if you want to shift quickly from understanding LLMs to building actual AI applications.

4. Natural Language Processing with Transformers – Lewis Tunstall, Leandro von Werra & Thomas Wolf

Where the previous book leans visual and practical, Natural Language Processing with Transformers goes deeper on engineering and professional deployment.

All three authors work at Hugging Face. You're essentially reading the official manual for the most popular open-source AI ecosystem today.

Content covers step-by-step training and deployment of BERT, GPT, T5, and others. You'll learn data preparation, training, fine-tuning, and evaluation workflows. Real-world examples span healthcare, finance, and multilingual NLP. The real concern here is that without practical hands-on experience, you might hit walls. This book assumes you're ready to actually implement things in production contexts. Essential reference material for ML engineers building real AI products with Hugging Face.

5. The LLM Engineering Handbook – Paul Iusztin & Maxime Labonne

Training a model? That's only half the battle. The harder part: turning that model into a stable product serving thousands or millions of users. That's what The LLM Engineering Handbook addresses.

Unlike training-focused books, this is a practitioner's guide for deploying AI systems. It covers the entire LLM lifecycle from initial research through production operations.

Topics include prompt optimization, function calling, tool use, advanced RAG architectures, and large-scale deployment strategies. All presented from a practical, real-world angle. This book is perfect for developers moving beyond simple API calls toward building scalable, reliable AI applications. If you're facing actual production constraints, this is your reference.

Which Book Should You Start With?

These five books form a complete learning progression from foundational theory to practical deployment.

If you're new to the field, start with The Hundred-Page Language Models Book or Hands-On Large Language Models. Both build intuition quickly and give you the big picture. Once foundations feel solid, Build a Large Language Model (From Scratch) teaches you how Transformers actually work under the hood. Next, Natural Language Processing with Transformers gets you productive with industry-standard Hugging Face workflows. Finally, when you're shipping real AI products, The LLM Engineering Handbook becomes your guide for optimization and large-scale operations.

In AI, the people who stand out aren't those who know the most prompts. They're the ones who understand what's actually happening inside the model behind those prompts. Wherever you are in your AI learning journey, at least one of these five deserves a spot on your shelf.


Description: Build real LLM expertise with this curated reading list—from foundational theory to production deployment.

Related Articles

Copyright © 2016 QTitHow All Rights Reserved