Search to Decision: A Journey Likely to End in a One Star Hotel
July 10, 2026
Another dinobaby post. No AI unless it is an image. This dinobaby is not Grandma Moses, just Grandpa Arnold.
I read a pretty wild and wooly essay intended for top dogs in organizations. My concern is that some of these deciders will fall for the razzle dazzle and end up in a bit of a swamp. The essay is “AI Knowledge Management Moves from Search Tool to Enterprise Decision Layer.” Yeah, okay. Enterprise search is not exactly a smooth running Toyota RAV in most organizations. Some information is not findable. Usually there are good reasons for the voids. (Drop into a pharma company and see if you can find info about a current clinical trial.)
Okay, Midjourney. Sort of disappointing.
The write up begins with a “typical” and I assume compelling example of a real life situation in the ideal corporate entity in the United States. Here’s the use case:
This demand is especially strong in organizations where information is spread across multiple systems. A salesperson may need product guidance from a knowledge base, a support agent may need context from past tickets and a product manager may be looking for insights from customer conversations. When that information lives in separate places, employees often spend time searching for answers or end up making decisions with only part of the picture. AI-powered platforms aim to reduce that friction. They can summarize long records, suggest relevant content and answer questions based on approved sources. The strongest systems also show where the answer came from, which helps users judge whether the information is current and reliable.
On the surface, the straw man seems reasonable. Let’s consider it from three angles.
First, the information required does not appear to be related to a law suit or the shroud of legal discovery. The information is not part of a government project operating under rules for classified information. The information does not seem to be that which is in emails, chats, or files on a computing device of an employee working remotely or on a device used by a contractor from a third party performing work for the organization. I am not sure if the information needed to answer certain questions is likely to be a system given indiscriminate access to the content in an organization.
Second, the old IDC chestnut that employees spend lots of time searching for information. Okay, but based on research in which I was involved at a blue chip consulting firm, employees find information this way: [a] A quick Google search, [b] Ask someone, [c] flip through local information on a laptop, a pile of folders on a credenza, etc. The searching angle does not hold up when employee work practices are observed, documented, and analyzed. A bonus insight: The closer one is to the top of the management hierarchy, asking and making a judgment call appear as a favorite method among a majority of senior managers.
Third, AI systems can output answers and suggestions based on approved sources. Okay? What is the time and cost to approve sources? How does an organization bumbling from one opportunity or crisis approve new sources, get them into the training set, and benefit from the flow of “new” information? The answer is, “Most AI systems can but don’t?” Why? How about cost and human fiddling around time? Based on the research I have done into information retrieval over the years, talk about fresh data available in real time is baloney. When an employee cannot locate the PowerPoint the sales person cooked up seal a deal confirmed in an email sent via Yahoo, that employee tries to “get in touch.” Yeah, good luck with that in today’s work environment.
Fourth, the user — that is, the employee who is fully informed, intelligent, and attentive k— will judge whether the information is current and reliable. What craziness is this? No employee knows if the information output is current, complete, and accurate. The painful truth is that people perceive the computer as being correct. This means that employees just use what’s output.
As you can tell, this write up is a marketing confection.
Here’s the conclusion to the write up:
AI-powered knowledge management is becoming more than enterprise search with a new interface. It is becoming a decision layer that connects people to usable institutional knowledge. The companies that succeed will be those that combine AI capability with governance discipline and practical workplace integration.
This passage contains a small nugget or uranium ore; to wit, “AI powered knowledge management is becoming more than enterprise search with a new interface.” Yes, I agree. It is going to become the glittering chaff of marketing pitches in the balance of 2026 and into 2027. Finding information is hard. Traditional enterprise search said, “No, it’s not.” Well, after the implosion of enterprise search vendors, licensees learned that search vendors were blowing smoke. Now the cycle is going to repeat. AI is a finding utility. Enterprise search is hard. AI is unlikely to make search better, faster, or cheaper, but it will definitely hallucinate and lead to some interesting decisions. Maybe companies should stick with paper, folders, and filing cabinets in separate organizational units. That one, one might know whom to ask for an answer. You may not get it, but at least you were close to a source.
Stephen E Arnold, July 10, 2026
Enterprise Search: A Story That Repeats Again and Again
July 2, 2026
Another dinobaby post. No AI unless it is an image. This dinobaby is not Grandma Moses, just Grandpa Arnold.
I spent some time looking to see if I have hard copies of the books I have written in my long, non-illustrious career. I have copies on the Enterprise Search Report, published more than a decade ago by an outfit of which I have lost track. The ESR was a best-seller. (That means I actually made a profit over the three editions of the 300 page document. Amazing!)
I included some discussion of Elastic and its predecessor Compass. Shay Banon was a person with a vision: Open source search and retrieval was the way to meet organization’s need for finding information it already possessed. The Elastic effort did reasonably well and spawned revenue via professional services, a pivot to cyber security via log file analysis, and a good marketing approach relying on training sessions in places like the Big Apple.

Does the coach’s advice make sense? Probably not. But coaches have to come up with something that does not create more problems for himself, his organization, and his gifted professionals. MidJourney, good enough.
Then two things happened, and I am not sure too many people noticed. I did, but after three editions of the ESR, I was eager to do work in other sectors. I kept my eye on a couple of outfits, and Elastic was one of these.
So what happened?
An outfit called Lucid Imagination later renamed to Lucid Works grabbed Lucene and Solr. The idea was that Lucid Words was a better commercial search and retrieval system. It’s still around, but the key event is not immediately clear (I wanted to write “lucid” but I didn’t.) Amazon hired a couple of go-getters from Lucid Works. That cross pollination of information was the seed that germinated into what I think of as the Amazon bulldozer driving slowly over Elastic. The result is that Amazon now offers search in flavors and one of those, in my opinion, is sure similar to Elastic’s approach.
The second event was the mostly accidental and quite messy spilling of AI technology into organizational information retrieval. Elastic jumped into that game, but it appears that people like using hallucinating artificial intelligence more than key word queries. I think that Elastic like Apple and Telegram are now suffering from the firm’s AI strategy or imprecision of its AI tactics.
If I were to write another Enterprise Search Report which I am not going to do, I would not include companies like Elastic or some of the firms that Gartner Group is presenting as enterprise search leaders. It’s time to put down those stone axes with AI trimmings. Just a thought from a dinobaby, of course.
Why did I take this quick trip down memory lane? I read “CEO Ash Kulkarni’s Organizational Announcement to Elastic Employees.” The write up said on June 24, 2026:
The industry is changing. Advances in AI, automation, and technology are reshaping how work gets done, and we’re changing with them. Customer expectations are increasing and evolving faster than ever before, and this requires us to move faster and operate leaner than we have before. We provide the data store and context engine for the technology that is changing the world, and we help observe and protect the infrastructure that runs it. We’re in a great position to lead here. To do it, we’re shifting our pace of innovation, simplifying how we operate, and investing in new skills. That’s what this reorganization is for: a simpler structure, with fewer layers, less complexity, and less friction.
My translation: “We’re need to cut costs. Hire more sales people who can close deals. Regroup to come up with a winning AI strategy. You are the first batch to be able to run a taco truck. More of you will have that opportunity in the next few months.”
I don’t think AI is particularly reliable unless the smart part is trained on quite specific content and then continuously updated to avoid the wonderful Bayesian drift like that suffered in the early Autonomy neural linguistic programming methods from the late 1990s. Drift is part of the game, and that’s why expensive training and retraining are part of the “fixing hallucination” function of the Googley transformer method.
Hybrid AI and traditional search models are unlikely to reduce the “cost” of finding information in an organization. Why? Organizations are silos of data. The legal department, particularly when engaged in litigation, is a little silo and sharing some information can have outsized downsides. The pharma and semiconductor R&D outfits use silos too. Powerful people in organizations expect silos. Index their mobile text messages and join the Elastic folks running a taco truck. Some government contracts require silos. Screw up government processes and you too can experience what I refer to as the “Anthropic wake up call.” Plus, organizations — contrary to popular belief — are busy outputting data and information, usually in a quite uncontrolled way.
The information at your fingertips pitch for enterprise search usually crashes into unseen shoals or get blown into rocks by business conditions.
Several observations:
- Enterprise search is in 2026 a difficult service to deliver, keep users happy, and costs low.
- Enterprise search without AI is expensive to pull off. That’s why most of the outfits I profiled in the three editions of the ESR are gone, morphed into some else, or embedded in some other enterprise service like document conversion.
- Enterprise search with AI is super difficult and potentially even more costly. US companies may not be comfortable loading their data into Qwen and hoping for the best in order to reduce token costs.
Net net: Elastic has had a good run. The big question is, “Can it handle the next leg of the enterprise software race ?”
Stephen E Arnold, July 2, 2026
ZSearch: Quite a Marketing Pitch
February 11, 2026
Another dinobaby post. No AI unless it is an image. This dinobaby is not Grandma Moses, just Grandpa Arnold.
I read “AI-Powered Enterprise Search: How ZSearch Redefines Organizational Knowledge Discovery.” The author is “Blitz.” Here is Blitz:

Blitz is a guest posts agency. The agency like categorical affirmatives. I counted them. There are 14 of them in the 860 word write up about ZSearch. A categorical affirmative is a word like all, every, always, and only. But that’s not all, I stuff the full text into one of my handy dandy word analysis tools and learned that 251 words qualified as marketing jargon and jingoisms. Yep, one third of the write up is puffery. That means that the host for the shaped content is pushing squishy information to an adoring group of AI content scrapers. And who is the “host”? It is something called Nerdbot.com. I don’t know much about this entity, and I assume they have the reader’s interest front and center.
What is ZSearch? From my point of view, a search and retrieval company with AI. Its marketing department, I would guess, has access to one of the online AI systems. The jargon density is first class. One would have to index and parse every output from Autonomy, Endeca, Fast Search & Transfer, and the other outfits that were trying to generate excitement for a search system for the enterprise.
Guess what? The marketing collateral was generally better than the performance of the enterprise search engine.
I think the optimal way to describe what ZSearch does to differentiate itself is to look at the jargon. Here’s an alphabetical list of the terms I extracted. Just scan the list and you will get a reasonable idea of what a customer can expect:
AI-Assisted Project Workspaces
AI-Driven / AI-Powered
Centralized Search Experience
Cloud Platforms
Compliance Standards
Context-Aware
Continuous Indexing
Conversational Search
Deterministic Outcomes
Distributed Data Sources
End-to-End Source Traceability
Enterprise-Grade
Hybrid Retrieval Model
Information Retrieval
Intelligent Enterprise Search
Knowledge Discovery
Knowledge Management
Metadata Extraction
Natural Language Processing (NLP)
Organizational Silos
Real-Time Synchronization
Scalable Architecture
Semantic Intelligence / Semantic Understanding
Traceability
Unified Access
User Intent
Workflows
After reading the write up, I had some questions; for example, What’s the security approach? How does one update the “index” to handle new terms or bound phrases? What’s the cost? What is the latency between discovery of new content and its availability to a user?
Should you license ZSearch? That’s up to you. Navigate to this location: https://zbrain.ai/zsearch/. You will discover an agent store and much more. ZSearch is a unit of ZBrain. The Web site doesn’t name the president or chief technical officer. It does not point out that ZBrain in Atlanta was acquired first by LeewayHertz. Then the Hackett Group purchased LeewayHertz. As far as I know, ZBrain and its ZSearch are part of the The Hackett Group (NASDAQ: HCKT). Its share price, according finance.google.com on February 7, 2026, appears to be $16.17 per share.
Enterprise search is tricky. Superlatives are easy to put in a marketing write up. They are tough to deliver in other settings.
Stephen E Arnold, February 11, 2026
File Conversion. No Problem. No Kidding?
December 10, 2025
Another short essay from a real and still-alive dinobaby. If you see an image, we used AI. The dinobaby is not an artist like Grandma Moses.
Every few months, I get a question about file conversion. The questions are predictable. Here’s a selection from my collection:
- “We have data about chemical structures. How can we convert these to for AI processing?”
- “We have back up files in Fastback encrypted format. How do we decrypt these and get the data into our AI system?”
- “We have some old back up tapes from our Burroughs’ machines?”
- “We have PDFs. Some were created when Adobe first rolled out Acrobat and some generated by different third-party PDF printing solutions. How can we convert these so our AI can provide our employees with access?”
The answer to each of these questions from the new players in AI-search system is, “No problem.” I hate to rain on these marketers’ assertions, but these are typical problems large, established organizations have moving content from a legacy system into a BAIT (big AI tech) based findability solution. There are technical challenges. There are cost challenges. There are efficiency challenges. That’s right. Challenges, and in my long career in electronic content processing, these hurdles still remain. But I am an aged dinobaby. Why believe me? Hire a Gartner-type of expert to tell you what you want to hear. Have fun with that solution, please.

Thanks, Venice.ai. Close enough for horse shoes, the high-water mark today I believe.
Venture Beat is one of my go-to sources for timely content marketing. On November 14, 2025, the venerable firm published “Databricks: PDF Parsing for Agentic AI Is Still Unsolved. New Tool Replaces Multi-Service Pipelines with a Single Function.” The write up makes clear that I am 100 percent dead wrong about processing PDF files with their weird handling of tables, charts, graphs, graphic ornaments, and dense financial data.
The write up explains how really off base I am; for example, the Databricks Agent Bricks Platform. It cracks the AI parsing problem. I learned from the Venture Beat write up identifies what the DABP does with PDF information:
1 “Tables preserved exactly as they appear, including merged cells and nested structures
2 Figures and diagrams with AI-generated captions and descriptions
3 Spatial metadata and bounding boxes for precise element location
4 Optional image outputs for multimodal search applications”
Once the PDFs have been processed by DABP, the outputs can be used in a number of ways. I assume these are advanced, stable, and efficient as the name “databrick” metaphorically suggests:
1 Spark declarative pipelines
2 Unity catalog (I don’t know what this means)
3 Vector search (yep, search and retrieval)
4 AI function chaining (yep, bots)
5 Multi-agent supervisor (yep, command and control).
The write up concludes with this statement:
The Databricks approach sheds new light on an issue that many might have considered to be a solved problem. It challenges existing expectations with a new architecture that could benefit multiple types of workflows. However, this is a platform-specific capability that requires careful evaluation for organizations not already using Databricks. For technical decision-makers evaluating AI agent platforms, the key takeaway is that document intelligence is shifting from a specialized external service to an integrated platform capability.
Net net: What is novel in that chemical structure? What about that guy who retired in 2002 who kept a pile of Fastback floppies with his research into in Trinitrotoluene variants? Yep, content processing is not problem except the data on those back up tapes cranked out by that old Burroughs’ MFSOLT utility, but with the new AI approaches, who needs layers of contractors and conversion utilities. Just say, “Not a problem.” Everything is easy for a market collateral writer.
Stephen E Arnold, December 10, 2025
Enterprise Search Is Back, Baby, or Is It Spelled Baiby
October 28, 2025
This essay is the work of a dumb dinobaby. No smart software required.
I want to be objective, but I admit I laughed. The spark for my jocularity was the marketing content document called “Claude AI Integrates with Microsoft 365 and Launches Enterprise Search for Business Teams.” But the subtitle tickled by aged ribs:
Anthropic has rolled out two new enterprise features for Claude, integration with Microsoft 365 and a unified search tool designed to connect organizational data across platforms.
There is one phrase to which I shall return, but let’s look at what the document presents as actual factual, ready to roll, enterprise ready software plus assorted AI magic dust.

Thanks, Venice.ai. Good enough.
I noted this statement about Microsoft / Anthropic or maybe Anthropic / Microsoft:
“Introducing two new features: Claude now connects to Microsoft 365 and offers enterprise search,” the company wrote on LinkedIn. “You can connect Claude to SharePoint, OneDrive, Outlook and Teams to search documents, analyze email threads, and review meeting summaries directly in conversation.” Anthropic, known for developing AI models focused on reliability and alignment, said both features are available immediately for Claude Team and Enterprise customers.
The Holy Grail is herewith ready for licensees to use to guzzle knowledge.
But what if the organization’s information is partitioned; for example, the legal department has confidential documents or is engaged in litigation and discovery is underway? What if the organization is in the pharmaceutical business, and the work is secret with a bit of medical trials activity underway. There are interview notes, laboratory data, and photographs of results? What if the organization lands a government contract with the Department of War and has to split off staff quickly as they transition from “regular” work to that which is conducted under quite specific rules and requirements? There are other questions as well; for example, what about those digitized schematics, the vendor information, and the data from digital cameras and work monitoring software?
I noted this statement as well:
Anthropic said the capability “brings your company’s knowledge into one place, using a dedicated shared project.” The system also includes custom prompts to refine searches and improve response accuracy. The company emphasized use cases such as onboarding new team members, identifying experts across departments, and analyzing feedback patterns to guide strategy and decision-making. The Microsoft 365 connector and enterprise search are now live for all Claude Team and Enterprise customers. Organization administrators can enable access and configure connected data sources.
My reaction, “This is 1990s enterprise search wearing a sweatshirt with a robot and AI on the front.”
Is this possible? Sure. The difficulty is that when employees interact with this type of system, interesting actions take place. One of the most common is, “This is not the document I wanted.” Often an employee will say, “This is not the PowerPoint I signed off for use at the conference.” Others may say, “Did you know that those salary schedules are in an Excel file with the documents about the employee picnic?”
Now let’s look at the phrase I thought interesting enough to discuss it in a separate paragraph. Here’s the phrase: “organizational data across platforms.” This evokes the idea that somewhere in the company is a cloud service containing corporate data. The Anthropic or Microsoft system will spider that content, process it, and make it findable. The hitch in the git along is that the other platforms may not embrace Microsoft security methods. Further the data may require a specific application to access those data. The cycle time between original indexing of the other platforms may be out of sync with the most recent data on those other platforms. In a magic world like the Fast Search & Transfer type environment which Microsoft purchased in 2008, the real world caused the magic carpet to lose altitude.
Now we have magic carpet 2025. How well with the painful realities of resources, security, cost, optimization, and infrastructure make it difficult for the magic carpet to keep flying? Marketing collateral pitches are easy. Delivering AI-enabled search to live up to the magic is slightly more difficult and, in my experience, shockingly expensive.
Stephen E Arnold, October 28, 2025
Same Old Search Problem, Same Old Search Solution
September 25, 2025
A problem as old as time is finding information within an organization. A good company organizes their information in paper and digital files, but most don’t do this. Digital information is arguably harder to find because you never know what hard drive or utility disc to search through. Apparently BlueDocs, via WRAL News, found a solution to this issue: “BlueDocs Unveils Revolutionary AI Global Search Feature, Transforming How Organizations Access Internal Documentation Software.”
The press release about BlueDocs, an AI global documentation software platform, opens with the usual industry and revolutionary jargon. Blah. Blah. Blah.
They have a special sauce:
“Unlike traditional search solutions that operate within platform boundaries, AI Global Search leverages advanced artificial intelligence to understand context, intent, and relationships across disparate knowledge sources. Users can now execute a single search query to simultaneously explore BlueDocs content, Google Workspace files, Microsoft 365 documents, and integrated third-party platforms.”
It delivers special results:
“ ‘AI Global Search has fundamentally changed how our team accesses information,’ said one Beta Customer. ‘What used to require checking five different platforms now happens with a single search. It’s particularly transformative for onboarding new team members who previously needed training on multiple systems just to find basic information.’”
Does this lingo sound like every other enterprise search solution’s marketing collateral? If BlueDocs delivers an easily programmable, out-of-the-box solution that interfaces across all platforms and returns usable results: EXCELLENT. If it needs extra tech support at a very special low price and custom engineering, the similarity with enterprise search of yore is back again.
Whitney Grace, September 25, 2025
What Happens When Content Management Morphs into AI? A Jargon Blast
September 16, 2025
Sadly I am a dinobaby and too old and stupid to use smart software to create really wonderful short blog posts.
I did a small project for a killer outfit in Cleveland. The BMW-driving owner of the operation talked about CxO this and CxO that. The jargon meant that “x” was a placeholder for titles like “Chief People Officer” or “Chief Relationship Officer” or some similar GenX concept.
I suppose I have a built in force shield to some business jargon, but I did turn off my blocker to read CxO Today’s marketing article titled helpfully “Gartner: Optimize Enterprise Search to Equip AI Assistants and Agents.” I was puzzled by the advertising essay, but then I realized that almost anything goes in today’s world of sell stuff by using jargon.
The write up is by an “expert” who used to work in the content management field. I must admit that I have zero idea what content management means. Like knowledge management, the blending of an undefined noun with the word “management” creates jargon that mesmerizes certain types of “leadership” or “deciders.”
The article (ad in essay form) is chock full of interesting concepts and words. The intent is to cause a “leadership” or “decider” to “reach out” for the consulting firm Gartner and buy reports or sit-downs with “experts.”
I noticed the term “enterprise search” in the title. What is “enterprise search” other than the foundation for the HP Autonomy dust up and the FAST Search & Transfer legal hassle? Most organizations struggle to find information that someone knows exists within an organization. “Leadership” decrees that “enterprise search” must be upgraded, improved, or installed. Today one can download an open source search system, ring up a cloud service offering remote indexing and search of “content,” or tap one of the super-well-funded newcomers like Glean or other AI-enabled search and retrieval systems.
Here’s what the write up advertorial says:
The advent of semantic search through vectorization and generative AI has revolutionized the way information is retrieved and synthesized. Search is no longer just an experience. It powers the experience by augmenting AI assistants. With RAG-based AI assistants and agents, relevant information fragments can be retrieved and resynthesized into new insights, whether interactively or proactively. However, the synthesis of accurate information depends largely on retrieving relevant data from multiple repositories. These repositories and the data they contain are rarely managed to support retrieval and synthesis beyond their primary application.
My translation of this jargon blast is that content proliferation is taking place and AI may be able to help “leadership” or a regular employee find the information needed to complete work. I mean who doesn’t want “RAG-based AI assistants” when trying to find a purchase order or to check the last quality report about a part that is failing 75 percent of the time for a big customer?
The fix is to embrace “touchpoints.” The write up says:
Multiple touchpoints and therefore multiple search services mean overlap in terms of indexes and usage. This results in unnecessary costs. These costs are both direct, such as licenses, subscriptions, compute and storage, and indirect, such as staff time spent on maintaining search services, incorrect decisions due to inaccurate information, and missed opportunities from lack of information. Additionally, relying on diverse technologies and configurations means that query evaluations vary, requiring different skills and expertise for maintenance and optimization.
To remediate this problem — that is, to deliver a useful enterprise search and retrieval system — the organization needs to:
aim for optimum touchpoints to information provided through maximum applications with minimum services. The ideal scenario is a single underlying service catering to all touchpoints, whether delivered as applications or in applications. However, this is often impractical due to the vast number of applications from numerous vendors… so
hire Gartner to figure out who is responsible for what, reduce the number of search vendors, and cut costs “by rationalizing the underlying search and synthesis services and associated technologies.”
In short, start over with enterprise search.
Several observations:
- Enterprise search is arguably more difficult than some other enterprise information problems. There are very good reasons for this, and they boil down to the nature of what employees need to do a job or complete a task
- AI is not going to solve the problem because these “wrappers” will reflect the problems in the content pools to which the systems have access
- Cost cutting is difficult because those given the job to analyze the “costs” of search discover that certain costs cannot be eliminated; therefore, their attendant licensing and support fees continue to become “pay now” invoices.
What do I make of this advertorial or content marketing item in CxO Today. First, I think calling it “news” is problematic. The write up is a bundle of jargon presented as a sales pitch. Second, the information in the marketing collateral is jargon and provides zero concrete information. And, third, the problem of enterprise search is in most organizational situations is usually a compromise forced on the organization because of work processes, legal snarls, secret government projects, corporate paranoia, and general turf battles inside the outfit itself.
The “fix” is not a study. The “fix” is not a search appliance as Google discovered. The “fix” is not smart software. If you want an answer that won’t work, I can identify whom not to call.
Stephen E Arnold, September 19, 2025
Glean Goes Beyond Search: Have Xooglers Done What Google Could Not Do?
August 12, 2025
This blog post is the work of an authentic dinobaby. Sorry. No smart software can help this reptilian thinker.
I read an interesting online essay titled “Glean’s $4.5B Business Model: How Ex-Googlers Built the Enterprise Search That Actually Works.” Enterprise search has been what one might call a Holy Grail application. Many have tried to locate the Holy Grail. Most have failed.
Have a small group of Xooglers (former Google employees) located the Holy Grail and been able to convert its power into satisfied customers? The essay, which reminded me of an MBA write up, argues that the outfit doing business as Glean has done it. The firm has found the Holy Grail, melted it down, and turned it into an endless stream of cash.
Does this sound a bit like the marketing pitch of Autonomy, Fast Search & Transfer, and even Google itself with its descriptions of its deeply wacky yellow servers? For me, Glean has done its marketing homework. The evidence is plumped and oiled for this essay about its business model. But what about search? Yeah, well, the focus of the marketing piece is the business model. Let’s go with what is in front of me. Search remains a bit of a challenge, particularly in corporations, government agencies, and pharmaceutical-type outfits where secrecy is a bit part of that type of organization’s way of life.
What is the Glean business model? It is VTDF. Here’s an illustration:
Does this visual look like blue chip consulting art? Is VTDF blue chip speak? Yes. And yes. For those not familiar with the lingo here’s a snapshot of the Glean business model:
- Value: Focuses on how the company creates and delivers core value to customers, such as solving specific problems
- Technology: Refers to the underlying tech innovations that allow “search” to deliver what employees need to do their jobs
- Distribution: Involves strategies for marketing, delivery, and reaching users
- Finance: Covers revenue models, cash flow management, and financial sustainability. Traditionally this has been the weak spot for the big-time enterprise search plays.
The essay explains in dot points that Glean is a “knowledge liberator.” I am not sure how that will fly in some pharma-type outfits or government agencies in which Palantir is roosting.
Once Glean’s “system” is installed, here’s what happens (allegedly):
- Single search box for everything
- Natural language queries
- Answers, not just documents
- Context awareness across apps
- Personalized to user permissions
- New employees productive in days.
I want to take a moment to comment on each of these payoffs or upsides.
First, a single search box for everything is going to present a bit of a challenge in several important use cases. Consider a company with an inventory control system, vendor evaluations, and a computer aid design and database of specifications. The single search box is going to return what for a specific part? Some users will want to know how many are in stock. Others will want to know the vendor who made the part in a specific batch because it is failing in use. Some will want to know what the part looks like? The fix for this type of search problem has been to figure out how to match the employee’s role with the filters applied that that user’s query. In the last 60 years, that approach sort of worked, but it was and still is incredibly difficult to keep lined up with employee roles, assorted permissions, and the way the information is presented to the person running the query. The quality issue may require stress analysis data and access to the lawsuit the annoyed customer has just filed. I am unsure how the Xooglers have solved this type of search task.
Second, the NLP approach is great but it is early 2000s. The many efforts, including DR-LINK to which my team contributed some inputs, were not particularly home run efforts. The reason has to do with the language skills of the users. Organizations hire people who may be really good at synthesizing synthetics but not so good at explaining what the new molecule does. If the lab crew dies, the answer does not require words. Querying for the “new” is tough, since labs doing secret research do not share their data. Even company officers have a tough time getting an answer. When a search system requires the researcher to input a query, that scientist may want to draw a chemical structure or input a query like this “C8N8O16.” Easy enough if the indexing system has access to the classified research in some companies. But the NLP problem is what is called “prompt engineering.” Most humans are just not very good at expressing what they need in the way of information. So modern systems try to help out the searcher. The reason Google search sucks is that the engineers have figured out how to deliver an answer that is good enough. For C8N8O16 close enough for horseshoes might be problematic.
Third, answer are what people want. The “if” statement becomes the issue. If the user knows a correct answer or just accepts what the system outputs. If the user understands the output well enough to make an informed decision. If the system understood or predicted what the user wanted. If the content is in the search systems index. This is a lot of ifs. Most of these conditions occur with sufficient frequency to kill outfits that have sold an “enterprise search system”.
Fourth, the context awareness across apps means that the system can access content on proprietary systems within an organization and across third party systems which may or may not run on the organization’s servers. Most enterprise search systems create or have licensed filters to acquire content. However, keeping the filters alive and healthy with the churn in permissions, file tweaks, and assorted issues related to latency creating data gaps remain tricky.
Fifth, the idea of making certain content available only to those authorized to view those data is a very tricky business. Orchestrating permissions is, in theory, easy to automate. The reality in today’s organizations is the complicating factor. With distributed outfits, contractors, and employees who may be working for another country secretly add some excitement to accessing “information.” The reality in many organizations is that there are regular silos like the legal department keeping certain documents under lock and key to projects for three letter agencies. In the pharma game, knowing “who” is working on a project is often a dead give-away for what the secret project is. The company’s “people” officer may be in the dark. What about consultants? What information is available to them? The reality is that modern organizations have more silos than the corn fields around Canton, Illinois.
Sixth, no training is required. “Employees are productive in days” is the pitch. Maybe, maybe not. Like the glittering generality that employees spend 20 percent of their time searching, the data for this assertion was lacking when the “old” IDC, Sue Feldman, and her team cranked out an even larger number. If anything, search is a larger part of work today for many people. The reasons range from content management systems which cannot easily be indexed in real time to the senior vice president of sales who changes prices for a product at a trade show and tells only his contact in the accounting department. Others may not know for days or months that the apple cart has been tipped.
Glean saves time. That is the smart software pitch. I need to see some data from a statistically valid sample with a reasonable horizontal x axis. The reference to “all” is troublesome. It underscores the immature understanding of what “enterprise search” means to a licensee versus what the venture backed company can actually deliver. Fast Search found out that a certain newspaper in the UK was willing to sue for big bucks because of this marketing jingo.
I want to comment briefly about “Technology Architecture: Beyond Search.” Hey, isn’t that the name of my blog which has been pumping out information access related articles for 17 years? Yep, it is.
Okay, Glean apparently includes these technologies in their enterprise search quiver:
- Universal connectors. Note the word “universal.” Nope, very tough.
- A Knowledge graph. Think in terms of Maltego, an open source software. Sure as long as there is metadata. But those mobile workers and their use of cloud services and EE2E messaging services. Sounds great. Execution in a cost sensitive environment takes a bit of work.
- An AI understanding layer. Yep, smart software. (Google’s smart software tells its users that it is ashamed of its poor performance. OpenAI rolled out ChatGPT 5 and promptly reposted ChatGPT 4o because enough users complained. Deepseek may have links to a nation state unfriendly to the US. Mark Zuckerberg’s Llama is a very old llama. Perplexity is busy fighting with Cloudflare. Anthropic is working to put coders out to pasture. Amazon, Apple, Microsoft, and Telegram are in the bolt it on business. The idea that Glean can understand [a] different employee contexts, [b] the rapidly changing real time data in an organization like that PowerPoint on the senior VP’s laptop, and [c] the file formats that have a very persistent characteristic of changing because whoever is responsible for an update or the format itself makes an intentional or unintentional change. I just can’t accept this assertion.
- Works instantly which I interpret as “real time.” I wonder if Glean can handle changed content in a legacy Ironside system running on AS/400s. I would sure like to see that and work up the costs for that cute real time trick. By the way, years ago, I got paid by a non US government agency to identify and define the types of “real time” data it had to process. I think my team identified six types. Only one could be processed without massive resource investments to make the other four semi real. The final one was to gain access to the high-speed data about financial instrument pricing in Wall Street big dogs. That simply was not possible without resources and cartwheels. The reason? The government wanted to search for who was making real time trades in certain financial instruments. Yeah, good luck with that in a world where milliseconds require truly big money for gizmos to capture the data and the software to slap metadata on what is little more than a jet engine exhaust of zeros and ones, often encrypted in a way that would baffle some at certain three letter agencies. Remember: These are banks, not some home brew messaging service.
There are some other wild assertions in the write up. I am losing interest is addressing this first year business school “analysis.” The idea is that a company with 500 to 50,000 employees can use this ready-to-roll service is interesting. I don’t know of a single enterprise search company I have encountered since I wrestled with IBM STAIRS and the dorky IBM CICS system that has what seems to be a “one size fits all” service. The Google Search Appliance failed with its “one size fits all.” The dead bodies on the enterprise search trail is larger than the death toll on the Oregon Trail. I know from my lectures that few if any know what DELPHES’ system did. What about InQuire? And there is IBM WebFountain and Clever. What about Perfect Search? What about Surfray? What about Arikus, Convera, Dieselpoint, or Entopia?
The good news is that a free trial is available. The cost is about $30 per month per user. For an organization like the local outfit that sells hard hats and uses Ironside and AS/400s, that works out to 150 times $360 or $54,000. I know this company won’t buy. Why? The system in place is good enough. Spreadsheet fever is not the same as identifying prospects and making a solid benefit based argument.
That’s why free and open source solutions get some love. Then built in “good enough” solutions from Microsoft are darned popular. Finally, some eager beaver in the information technology department will say, “Let me put together a system using Hugging Face.”
Many companies and a number of quite intelligent people (including former Googlers) have tried to wrestle enterprise search to the ground. Good luck. Just make sure you have verifiable data and not the wild assertions about how much time spend searching or how much time an employee will save. Don’t believe anything about enterprise search that uses the words “all” or universal.”
Google said it was “universal search.” Yeah, why after decades of selling ads does the company provide so so search for the Web, Gmail, YouTube, and images. Just ask, “Why?” Search is a difficult challenge.
Glean this from my personal opinion essay: Search is difficult, and it has yet to be solved except for precisely defined use cases. Google experience or not, the task is out of reach at this time.
Stephen E Arnold, August 12, 2025
Ground Hog Day: Smart Enterprise Search
January 7, 2025
I am a dinobaby. I also wrote the Enterprise Search Report, 1st, 2nd, and 3rd editions. I wrote The New Landscape of Search. I wrote some other books. The publishers are long gone, and I am mostly forgotten in the world of information retrieval. Read this post, and you will learn why. Oh, no AI helped me out unless I come up with an art idea. I used Stable Diffusion for the rat, er, sorry, ground hog day creature.
I think it was 2002 when the owner of a publishing company asked me if I thought there was an interest in profiles of companies offering “enterprise search solutions.” I vaguely remember the person, and I will leave it up to you to locate a copy of the 400 page books I wrote about enterprise search.
The set up for the book was simple. I identified the companies which seemed to bid on government contracts for search, companies providing search and retrieval to organizations, and outfits which had contacted me to pitch their enterprise search systems before they were exiting stealth mode. By the time the first edition appeared in 2004, the companies in the ESR were flogging their products.
The ground hog effect is a version of the Yogi Berra “Déjà vu all over again” thing. Enterprise search is just out of reach now and maybe forever.
The enterprise search market imploded. It was there and then it wasn’t. Can you describe the features and functions of these enterprise search systems from the “golden age” of information retrieval:
- Innerprise
- InQuira
- iPhrase
- Lextek Onix
- MondoSearch
- Speed of Mind
- Stratify (formerly Purple Yogi)
The end of enterprise search coincided with large commercial enterprises figuring out that “search” in a complex organization was not one thing. The problem remains today. Lawyers in a Fortune 1000 company want one type of search. Marketers want another “flavor” of search. The accountants want a search that retrieves structured and unstructured data plus images of invoices. Chemists want chemical structure search. Senior managers want absolutely zero search of their personal and privileged data unless it is lawyers dealing with litigation. In short, each unit wants a highly particularized search and each user wants access to his or her data. Access controls are essential, and they are a hassle at a time when the notion of an access control list was like learning to bake bread following a recipe in Egyptian hieroglyphics.
These problems exist today and are complicated by podcasts, video, specialized file types for 3D printing, email, encrypted messaging, unencrypted messaging, and social media. No one has cracked the problem of a senior sales person who changes a PowerPoint deck to close a deal. Where is that particular PowerPoint? Few know and the sales person may have deleted the file changed minutes before the face to face pitch. This means that baloney like “all” the information in an organization is searchable is not just stupid; it is impossible.
The key events were the legal and financial hassles over Fast Search & Transfer. Microsoft bought the company in 2008 and that was the end of a reasonably capable technology platform and — believe it or not — a genuine alternative to Google Web search. A number of enterprise search companies sold out because the cost of keeping the technology current and actually running a high-grade sales and marketing program spelled financial doom. Examples include Exalead and Vivisimo, among others. Others just went out of business: Delphes (remember that one?). The kiss of death for the type of enterprise search emphasized in the ESR was the acquisition of Autonomy by Hewlett Packard. There was a roll up play underway by OpenText which has redefined itself as a smart software company with Fulcrum and BRS Search under its wing.
What replaced enterprise search when the dust settled in 2011? From my point of view it was Shay Banon’s Elastic search and retrieval system. One might argue that Lucid Works (né Lucid Imagination) was a player. That’s okay. I am, however, to go with Elastic because it offered a version as open source and a commercial version with options for on-going engineering support. For the commercial alternatives, I would say that Microsoft became the default provider. I don’t think SharePoint search “worked” very well, but it was available. Google’s Search Appliance appeared and disappeared. There was zero upside for the Google with a product that was “inefficient” at making a big profit for the firm. So, Microsoft it was. For some government agencies, there was Oracle.
Oracle acquired Endeca and focused on that computationally wild system’s ability to power eCommerce sites. Oracle paid about $1 billion for a system which used to be an enterprise search with consulting baked in. One could buy enterprise search from Oracle and get structured query language search, what Oracle called “secure enterprise search,” and may a dollop of Triple Hop and some other search systems the company absorbed before the end of the enterprise search era. IBM talked about search but the last time I drove by IBM Government systems in Gaithersburg, Maryland, it like IBM search, had moved on. Yo, Watson.
Why did I make this dalliance on memory lane the boring introduction to a blog post? The answer is that I read “Are LLMs At Risk Of Going The Way Of Search? Expect A Duopoly.” This is a paywalled article, so you will have to pony up cash or go to a library. Here’s an abstract of the write up:
The evolution of LLMs (Large Language Models) will lead users to prefer one or two dominant models, similar to Google’s dominance in search.
Companies like Google and Meta are well-positioned to dominate generative AI due to their financial resources, massive user bases, and extensive data for training.
Enterprise use cases present a significant opportunity for specialized models.
Therefore, consumer search will become a monopoly or duopoly.
Let’s assume the Forbes analysis is accurate. Here’s what I think will happen:
First, the smart software train will slow and a number of repackagers will use what’s good enough; that is, cheap enough and keeps the client happy. Thus, a “golden age” of smart search will appear with outfits like Google, Meta, Microsoft, and a handful of others operating as utilities. The US government may standardize on Microsoft, but it will be partners who make the system meet the quite particular needs of a government entity.
Second, the trajectory of the “golden age” will end as it did for enterprise search. The costs and shortcomings become known. Years will pass, probably a decade, maybe less, until a “new” approach becomes feasible. The news will diffuse and then a seismic event will occur. For AI, it was the 2023 announcement that Microsoft and OpenAI would change how people used Microsoft products and services. This created the Google catch up and PR push. We are in the midst of this at the start of 2025.
Third, some of the problems associated with enterprise information and an employee’s finding exactly what he or she needs will be solved. However, not “all” of the problems will be solved. Why? The nature of information is that it is a bit like pushing mercury around. The task requires fresh thinking.
To sum up, the problem of search is an excellent illustration of the old Hegelian chestnut of Hegelian thesis, antithesis, and synthesis. This means the problem of search is unlikely to be “solved.” Humans want answers. Some humans want to verify answers which means that the data on the sales person’s laptop must be included. When the detail oriented human learns that the sales person’s data are missing, the end of the “search solution” has begun.
The question “Will one big company dominate?” The answer is, in my opinion, maybe in some use cases. Monopolies seem to be the natural state of social media, online advertising, and certain cloud services. For finding information, I don’t think the smart software will be able to deliver. Examples are likely to include [a] use cases in China and similar countries, [b] big multi-national organizations with information silos, [c] entities involved in two or more classified activities for a government, [d] high risk legal cases, and [e] activities related to innovation, trade secrets, and patents, among others.
The point is that search and retrieval remains an extraordinarily difficult problem to solve in many situations. LLMs contribute some useful functional options, but by themselves, these approaches are unlikely to avoid the reefs which sank the good ships Autonomy and Fast Search & Transfer, and dozens of others competing in the search space.
Maybe Yogi Berra did not say “Déjà vu all over again.” That’s okay. I will say it. Enterprise search is “Déjà vu all over again.”
Stephen E Arnold, January 7, 2025
The Hay Day of Search Has a Ground Hog Moment
December 19, 2024
This blog post is the work of an authentic dinobaby. No smart software was used.
I think it was 2002 or 2003 that I started writing the first of three editions of Enterprise Search Report. I am not sure what happened to the publisher who liked big, fat thick printed books. He has probably retired to an island paradise to ponder the crashing blue surf.
But it seems that the salad days of enterprise search are back. Elastic is touting semantics, smart software, and cyber goodness. IBM is making noises about “Watson” in numerous forms just gift wrapped with sparkly AI ice cream jimmies. There is a start up called Swirl. The HuggingFace site includes numerous references to finding and retrieving. And there is Glean.
I keep seeing references to Glean. When I saw a link to the content marketing piece “Glean’s Approach to Smarter Systems: AI, Inferencing and Enterprise Data,” I read it. I learned that the company did not want to be an AI outfit, a statement I am not sure how to interpret; nevertheless, the founder of Glean is quoted as saying:
“We didn’t actually set out to build an AI application. We were first solving the problem of people can’t find anything in their work lives. We built a search product and we were able to use inferencing as a core part of our overall product technology,” he said. “That has allowed us to build a much better search and question-and-answering product … we’re [now] able to answer their questions using all of their enterprise knowledge.”
And what happened to finding information? The company has moved into:
- Workflows
- Intelligent data discovery
- Problem solving
And the result is not finding information:
Glean enables enterprises to improve efficiency while maintaining control over their knowledge ecosystem.
Translation: Enterprise search.
The old language of search is gone, but it seems to me that “search” is now explained with loftier verbiage than that used by Fast Search & Transfer in a lecture delivered in Switzerland before the company imploded.
Is it now time for write the “Enterprise Knowledge Ecosystem Report”? Possibly for someone, but it’s Ground Hog time. I have been there and done that. Everyone wants search to work. New words and the same challenges. The hay is growing thick and fast.
Stephen E Arnold, December 19, 2024

