Microsoft Marketing Scores an Own Goal
June 25, 2026
Another dinobaby post. No AI unless it is an image. This dinobaby is not Grandma Moses, just Grandpa Arnold.
Absolutely the class action referenced in “Microsoft Shareholders Sue Over Allegedly Overhyped AI Performance and Hidden Cloud Slump” may be without merit. “Merit” is a slippery concept in today’s go go AI octagon. The issue seems to be related to marketing; that is, adding some sizzle to make that dried beef shingle worthy.

Thanks, MidJourney. Your art output reminded me of my exciting commutes on the 101. Rural Kentucky may not do the AI thing, but the traffic is definitely better than some AI systems’ outputs.
The write up says:
Microsoft is facing a class action lawsuit from shareholders accusing the company of defrauding investors by overhyping its AI products and hiding weaknesses in its cloud business. The complaint was filed in federal court in Seattle by the City of St. Clair Shores Police and Fire Retirement System, based in Michigan. Microsoft has dismissed the allegations as without merit and said it will defend itself in court. “Microsoft stands by the integrity of its public statements and will vigorously defend itself,” the company stated.
Okay, but a police and fire retirement system is likely to generate some media coverage in outlets from Police1 to the Telegram law enforcement channels. With Microsoft hoping to retain its grip on its US enterprise business, the Google is likely to perk up and monitor this legal matter as well.
The write up continues:
The shareholders allege that Microsoft portrayed its partnership with OpenAI as strong and stable despite ongoing fragmentation between the two companies. They also claim that Microsoft promoted Copilot adoption more aggressively than the results warranted. Additionally, they argue the company failed to adequately disclose or played down the financial commitments needed to build the data centers required for advanced AI models. The complaint further states that Microsoft did not properly disclose a decline in cloud revenue while continuing to invest billions in AI infrastructure without seeing proportionate returns.
What happens if one of the retirement system’s legal eagles stumbles upon the remarkable Ed Zitron? That individual has clawed his way (yeah, that’s a pun) to the top of the “AI is a problem” heap of pundits. What if that same group of legal eagles get in touch with Dr. Gary Marcus. Those two could add some spice to the allegations that big, management-starved Microsoft is arguably better at:
- Blowing its PR lead in AI to the estimable Google
- Floundering around in the kiddie pool with OpenAI and a number of enlistees as other companies deployed more useful and reliable smart software to help people with routine office tasks
- Pulling off one of those fancy moves that are the trademark of the tango champions Yanina Quiñones and Neri Piliu with the breath-taking “yes, it is Copilot and no, it is different” AI flourishes.
The write up notes:
The lawsuit is a formal legal challenge to Microsoft’s AI claims, which are among the more heavily promoted in the industry. If the case moves forward, discovery could reveal internal communications about Copilot adoption metrics, tensions in the OpenAI partnership, and infrastructure cost forecasts…. The case could become a reference point for similar lawsuits if other AI companies report disappointing results that fall short of investor expectations.
I don’t have a dog in this fight. Personally I am skeptical of the current crop of smart software systems. I do pay attention to quite specific applications of AI in certain policeware and intelware services. These are usually constrained and provide quite useful benefits to investigators working certain types of cases. But AI to design a treatment for my grand daughter’s illness? Sorry, none of the AI outfits have my confidence. I spent too much time analyzing the IBM Watson – Houston cancer thing.
Many interesting AI “actions” are underway. However, the combination of that living management case study in the Seattle area and a fire and law enforcement retirement system is different.
Net net: Hire Ed and Gary as expert witnesses.
Stephen E Arnold, June 25, 2026
The Search Engine Graveyard: A New Resident
May 5, 2026
Another dinobaby post. No AI unless it is an image. This dinobaby is not Grandma Moses, just Grandpa Arnold.
I was working for a search-and-retrieval company when AskJeeves.com became available in 1997. As it turned out, the natural language breakthrough that set AskJeeves apart from the other Web search engines was its question-answering angle. The firm at which I worked hired “content specialists.” From interviewing job seekers, I learned that AskJeeves’ approach was to can certain common questions. The answers to these questions would be updated. Some were automated like “What’s the weather in San Francisco?” but others required a human to craft a response. Other queries were passed to a search-and-retrieval system. Manual processes here are expensive. AskJeeves, therefore, bought “promising” companies for their indexing and content processing capabilities; for example, Jigsaw Technologies in 2000, Direct Hit Technologies in 2000 (specializing in search result ranking), and Teoma Technologies in 2001. AskJeeves tried repurposing its technology for customer service. But Google was maturing into the organization we all know today. In 2005, Barry Diller added AskJeeves to his collection of Internet properties. After the acquisition, Mr. Diller learned that Web search was a difficult and expensive business. The Ask.com service became a metasearch system, recycling search results from other Web indexing outfits in an effort to reduce costs.

Mashable has now reported that Ask.com is dead. “Every Great Search Must Come to an End” said:
Amid an overwhelming shift toward generative AI-powered search engines and a repositioning of AI agents as the future of web browsing, the loss of Ask.com feels like a true end of the early dot-com era. So long Jeeves, hello AI.
I want to add a bit of color to the demise of this Web search system.
My view is that smart software is indeed search-and-retrieval, just with bells and whistles. Systems like AskJeeves knew that handling queries from users was a tricky business. A certain percentage of queries were repetitive. These could be created and later cached. The acquisitions made clear that the original founders could not innovate in substantive ways. Garrett Gruener and David Warthen could recognize interesting technology and its applications. The acquisitions added some scope to the AskJeeves service, but financial realities sparked a sale to Barry Diller’s IAC in 2005. Web search became the province of deep-pocket entities like Google and Microsoft. These firms’ money came from reasonably solid revenue streams. Google sold ads and its pay-to-play model, and Microsoft licensed software. Without meaningful regulation, Google-type organizations trampled over companies like Lycos and All-the-Web, among others. .
This means that today, search-and-retrieval technology exists but has adopted a new vocabulary. The constants are the same: Expensive, complex, and expensive. Did I mention expensive?
The trajectory of AskJeeves is essentially the same for other search-and-retrieval enterprises: Rollout, technical enhancement, utility function, and disappearance or replacement by a spiffed-up version of the old stuff. If this sounds like the trajectory of artificial intelligence, I have made my point. One can apply this general pattern to Autonomy plc, Fast Search & Transfer, and dozens of search-and-retrieval systems that did not evolve into viable businesses. The technology may chug along in a content management system or may be used to perform a background activity, but the spotlight is not on old-school content search. Instead, attention is paid to smart software that requires massive infrastructure to do what humans did for AskJeeves. I would suggest that human-intermediated systems are more common than the marketers want to communicate. Therefore, AI is probably going to follow an AskJeeves type of fate over the next decade or two.?
Why do I suggest this? Here are my reasons based on my research while writing several books about search, including The New Landscape of Search, CyberOSINT: Next Generation Information Access, and The Enterprise Search Report 1st, 2nd, and 3rd editions, among others.
- Indexing can be automated, but one must know what words or phrases to use in the query in order to match certain content. A search in Bing, Google, or Yandex for “financial fraud” will not allow a teen to become a criminal in 10 minutes. Enter the term “carding,” and the game changes. Even today, software cannot replicate this “lingo knowledge.” Many tricks are used to try to know what the user really wants, but these fall short. The tricks like “field codes” themselves become because a person looking for information must know the code to get the chunked results..
- Content is fluid. Language is fluid. Search systems such as those used by Dialog’s or SDC perform best with static terminology. Scholars like static terminology. Indexing conventions try to cope with contextual issues; for example, does “terminal” mean “train station” or does it mean “mainframe peripheral”? The money pumped into smart software is trying to solve this basic problem for many user queries (or in new lingo, “user prompts”).
- The context of information is [a] volatile because today’s problem may not have existed yesterday and [b] situational; that is, every user operates within an “information ecosystem.” Outsiders have a tough time knowing what the characteristics of the ecosystem imply; for example, “loca” may mean one thing to a YouTube cruise personality and another thing to a person working in nuclear safety engineering. That’s why the efforts at personalization are becoming increasingly invasive. Ecosystem information is needed to provide somewhat useful outputs. What if that ecosystem is classified? Well, the big vendors don’t care. They will take what they can get because without it, the outputs are likely to be wrong or potentially quite problematic.
With the reality of change in these three facets of search-and-retrieval, it is appropriate to appreciate the efforts so many people have contributed to making “search” better. Too bad that most of these systems have failed and burned massive sums of money as they trail flames and smoke across the conference rooms in which revenue talks are held.
I have resisted writing about smart software. Everyone I meet is convinced that artificial intelligence is, by golly, the next big thing. Okay. I have other topics to research. I do want to remind readers that smart software is nothing more than search software wearing the latest designer jeans. That does not make it bad. I think the current skepticism about AI is a normal reaction to the discovery that hallucinations, high costs, and AI systems making decisions about health care, education, and judicial actions will present some problems going forward.
Remember. Search is difficult. Knowledge value requires verifiable facts and a foundation of generally accepted information. Without that, system outputs are useless and potentially harmful. Search gets traction because the systems so far developed don’t quite solve a user’s problem. Thus, search is a work in progress, and that progress is expensive. Mr. Diller pulled the plug.
I want to add a bit of color to the demise of this Web search system.
My view is that smart software is indeed search-and-retrieval just with bells and whistles. Systems like AskJeeves knew that handling queries from users was a tricky business. A certain percentage of queries were repetitive. These could be canned and latter cached. The acquisitions made clear that the original ideas and the original founders could not innovate in substantive ways. The founders, Garrett Gruener and David Warthen, could recognize interesting technology and its applications. The acquisitions added some scope to the AskJeeves service, but financial realities sparked a sale to in 2005. Web search became the province of deep pocket outfits like Google and Microsoft. These firms’ money came from reasonably solid revenue streams. Google sold ads or the pay-to-play model and Microsoft licensed software. Without meaningful regulation, Google-type outfits trampled over Lycos- and All-the-Web type outfits.
This means that search-and-retrieval today exists but it has adopted a new vocabulary. The constants are the same: Expensive, complex, and expensive. Did I mention expensive?
The trajectory of AskJeeves is essentially the same for other search-and-retrieval outfits: Roll out, technical enhancement, utility function, and disappearance or replacement by the old stuff spiffed up. If this sounds like the trajectory of artificial intelligence, I have made my point. One can apply this general trajectory to Autonomy plc, Fast Search & Transfer, and dozens of search-and-retrieval systems that have not evolved into viable businesses. The technology may chug along in a content management system or be used to perform a background activity. But the spotlight is not on old-school search-and-retrieval. The bright new manifestations of search and retrieval capture attention. Hint: smart software that requires massive infrastructure to do what humans did for AskJeeves. I would suggest that human-intermediated systems are more common than the marketers want to communicate. Therefore, AI is probably going to follow an AskJeeves type of trajectory over the next decade or two.
Why do I suggest this? Here are my reasons based on my research and writing of a number of books about search, including The New Landscape of Search, CyberOSINT: Next Generation Information Access, and The Enterprise Search Report 1st, 2nd, and 3rd editions, among others.
- Indexing can be automated but one has to know the words or phrases to use in the query in order to match certain content. Today one can navigate to Bing, Google, or Yandex and search “financial fraud.” The results will not allow a teen to become a criminal in 10 minutes. Enter the term “carding” and the game changes. Even today, software cannot replicate this “lingo knowledge.” Many tricks are used to try to know what the user really wants, but these fall short. The tricks themselves become problematic.
- Content is fluid. Language is fluid. Search-and-retrieval, whether old-school like Dialog Information’s or SDC’s approach, likes static terminology. Scholars like static terminology. Indexing conventions try to cope with contextual issues; for example, does “terminal” mean train station or does it mean “mainframe peripheral”? The money pumped into smart software is trying to solve this basic problem for many user queries or in new lingo “user prompts”.
- The context of information is [a] volatile because today’s problem may not have existed yesterday and [b] situational; that is, every user exists within an “information ecosystem.” Outsiders have a tough time knowing what the characteristics of the ecosystem mean; for example, “loca” may mean one thing to a YouTube cruise personality and another thing to a person working in nuclear safety engineering. That’s why the efforts at personalization are becoming increasingly invasive. Ecosystem information is needed to provide useful outputs. What if that ecosystem is classified? Well, the big vendors don’t care. They will take the information because without those data, the outputs are likely to be wrong or potentially quite problematic.
With the reality of change in these three facets of search-and-retrieval, one has to appreciate the efforts so many people have contributed to making “search” better. Too bad that most of these systems have failed and burned massive sums of money as they trail flames and smoke across the conference rooms in which revenue talks are held.
I have resisted writing about smart software. Everyone I meet is convinced that artificial intelligence is — by golly — the next big thing. Okay. I have other topics to research. I do want to remind anyone reading this short blog post that smart software is nothing more than search and retrieval wearing the latest designer jeans. That does not make it bad. I think the current skepticism about AI is a normal reaction to people discovering that hallucinations, high costs, and specter of AI systems making decisions about health care, education, and judicial actions is going to present some problems going forward.
Remember. Search and retrieval are difficult. Knowledge value requires verifiable facts and a foundation of generally accepted information. Without that system outputs are useless and potentially harmful. Search gets traction because the systems don’t quite solve the user’s problem. Thus, search is a work in progress, and that progress is expensive. Mr. Diller pulled the plug.
Stephen E Arnold, May 5, 2026
Problematic Smart Algorithms
December 12, 2023
This essay is the work of a dumb dinobaby. No smart software required.
We already know that AI is fundamentally biased if it is trained with bad or polluted data models. Most of these biases are unintentional due ignorance on the part of the developers, I.e. lack diversity or vetted information. In order to improve the quality of AI, developers are relying on educated humans to help shape the data models. Not all of the AI projects are looking to fix their polluted data and ZD Net says it’s going to be a huge problem: “Algorithms Soon Will Run Your Life-And Ruin It, If Trained Incorrectly.”
Our lives are saturated with technology that has incorporated AI. Everything from an application used on a smartphone to a digital assistant like Alexa or Siri uses AI. The article tells us about another type of biased data and it’s due to an ironic problem. The science team of Aparna Balagopalan, David Madras, David H. Yang, Dylan Hadfield-Menell, Gillian Hadfield, and Marzyeh Ghassemi worked worked on an AI project that studied how AI algorithms justified their predictions. The data model contained information from human respondents who provided different responses when asked to give descriptive or normative labels for data.
Normative data concentrates on hard facts while descriptive data focuses on value judgements. The team noticed the pattern so they conducted another experiment with four data sets to test different policies. The study asked the respondents to judge an apartment complex’s policy about aggressive dogs against images of canines with normative or descriptive tags. The results were astounding and scary:
"The descriptive labelers were asked to decide whether certain factual features were present or not – such as whether the dog was aggressive or unkempt. If the answer was "yes," then the rule was essentially violated — but the participants had no idea that this rule existed when weighing in and therefore weren’t aware that their answer would eject a hapless canine from the apartment.
Meanwhile, another group of normative labelers were told about the policy prohibiting aggressive dogs, and then asked to stand judgment on each image.
It turns out that humans are far less likely to label an object as a violation when aware of a rule and much more likely to register a dog as aggressive (albeit unknowingly ) when asked to label things descriptively.
The difference wasn’t by a small margin either. Descriptive labelers (those who didn’t know the apartment rule but were asked to weigh in on aggressiveness) had unwittingly condemned 20% more dogs to doggy jail than those who were asked if the same image of the pooch broke the apartment rule or not.”
The conclusion is that AI developers need to spread the word about this problem and find solutions. This could be another fear mongering tactic like the Y2K implosion. What happened with that? Nothing. Yes, this is a problem but it will probably be solved before society meets its end.
Whitney Grace, December 12, 2023
Word Problems Are Tricky for AI Language Models
October 27, 2022
If you have trouble with word problems, rest assured you are in good company. Machine-learning researchers have only recently made significant progress teaching algorithms the concept. IEEE Spectrum reports, “AI Language Models Are Struggling to ‘Get’ Math.” Writer Dan Garisto states:
“Until recently, language models regularly failed to solve even simple word problems, such as ‘Alice has five more balls than Bob, who has two balls after he gives four to Charlie. How many balls does Alice have?’ ‘When we say computers are very good at math, they’re very good at things that are quite specific,’ says Guy Gur-Ari, a machine-learning expert at Google. Computers are good at arithmetic—plugging numbers in and calculating is child’s play. But outside of formal structures, computers struggle. Solving word problems, or ‘quantitative reasoning,’ is deceptively tricky because it requires a robustness and rigor that many other problems don’t.”
Researchers threw a couple datasets with thousands of math problems at their language models. The students still failed spectacularly. After some tutoring, however, Google’s Minerva emerged as a star pupil, having achieved 78% accuracy. (Yes, the grading curve is considerable.) We learn:
“Minerva uses Google’s own language model, Pathways Language Model (PaLM), which is fine-tuned on scientific papers from the arXiv online preprint server and other sources with formatted math. Two other strategies helped Minerva. In ‘chain-of-thought prompting,’ Minerva was required to break down larger problems into more palatable chunks. The model also used majority voting—instead of being asked for one answer, it was asked to solve the problem 100 times. Of those answers, Minerva picked the most common answer.”
Not a practical approach for your average college student during an exam. Researchers are still not sure how much Minerva and her classmates understand about the answers they are giving, especially since the more problems they solve the fewer they get right. Garisto notes language models “can have strange, messy reasoning and still arrive at the right answer.” That is why human students are required to show their work, so perhaps this is not so different. More study is required, on the part of both researchers and their algorithms.
Cynthia Murrell, October 27, 2022
Smart Software and Textualists: Are You a Textualist?
June 13, 2022
Many thought it was simply a massive bad decision from an inexperienced judge. But there was more to it—it was a massive bad decision from an inexperienced textualist judge with an overreliance on big data. The Verge discusses “The Linguistics Search Engine that Overturned the Federal Mask Mandate.” Search is useful, but it must be accompanied by good judgment. When a lawsuit challenging the federal mask mandate came across her bench, federal judge Kathryn Mizelle turned to the letter of the law. Literally. Reporter Nicole Wetsman tells us:
“Mizelle took a textualist approach to the question — looking specifically at the meaning of the words in the law. But along with consulting dictionaries, she consulted a database of language, called a corpus, built by a Brigham Young University linguistics professor for other linguists. Pulling every example of the word ‘sanitation’ from 1930 to 1944, she concluded that ‘sanitation’ was used to describe actively making something clean — not as a way to keep something clean. So, she decided, masks aren’t actually ‘sanitation.’”
That is some fine hair splitting. The high-profile decision illustrates a trend in US courts that has been growing since 2018—basing legal decisions on large collections of texts meant for academic exploration. The article explains:
“A corpus is a vast database of written language that can include things like books, articles, speeches, and other texts, amounting to hundreds of millions of lines of text or more. Linguists usually use corpora for scholarly projects to break down how language is used and what words are used for. Linguists are concerned that judges aren’t actually trained well enough to use the tools properly. ‘It really worries me that naive judges would be spending their lunch hour doing quick-and-dirty searches of corpora, and getting data that is going to inform their opinion,’ says Mark Davies, the now-retired Brigham Young University linguistics professor who built both the Corpus of Contemporary American English and the Corpus of Historical American English. These two corpora have become the tools most commonly used by judges who favor legal corpus linguistics.”
Here is an example of how a lack of careful consideration while using the corpora can lead to a bad decision: the most frequent usage of a particular word (like “sanitation”) is not always the most commonly understood usage. Linguists emphasize the proper use of these databases requires skilled interpretation, a finesse a growing number of justices either do not possess or choose not to use. Such textualists apply a strictly literal interpretation to the words that make up a law, ignoring both the intent of lawmakers and legislative history. This approach means judges can avoid having to think too deeply or give reasons on the merits for their interpretations. Why, one might ask, should we have justices at all when we could just ask a database? Perhaps we are headed that way. We suppose it would save a lot of tax dollars.
See the article for more on legal corpora and how judges use them, textualism, and the problems with this simplified approach. If judges won’t respect the opinion of the very authors of the corpora on how they should and should not be used, where does that leave us?
Cynthia Murrell, June 13, 2022
Deepset: Following the Trail of DR LINK, Fast Search and Transfer, and Other Intrepid Enterprise Search Vendors
April 29, 2022
I noted a Yahooooo! news story called “Deepset Raises $14M to Help Companies Build NLP Apps.” To me the headline could mean:
Customization is our business and services revenue our monetization model
Precursor enterprise search vendors tried to get gullible prospects to believe a company could install software and employees could locate the information needed to answer a business question. STAIRS III, Personal Library Software / SMART, and the outfit with forward truncation (InQuire) among others were there to deliver.
Then reality happened. Autonomy and Verity upped the ante with assorted claims. The Golden Age of Enterprise Search was poking its rosy fingers through the cloud of darkness related to finding an answer.
Quite a ride: The buzzwords sawed through the doubt and outfits like Delphis, Entopia, Inference, and many others embraced variations on the smart software theme. Excursions into asking the system a question to get an answer gained steam. Remember the hand crafted AskJeeves or the mind boggling DR LINK; that was, document retrieval via linguistic knowledge.
Today there are many choices for enterprise search: Free Elastic, Algolia, Funnelback now the delightfully named Squiz, Fabasoft Mindbreeze, and, of course, many, many more.
Now we have Deepset, “the startup behind the open source NLP framework Haystack, not to be confused with Matt Dunie’s memorable “haystack with needles” metaphor, the intelware company Haystack, or a basic piles of dead grass.
The article states:
CEO Milos Rusic co-founded Deepset with Malte Pietsch and Timo Möller in 2018. Pietsch and Möller — who have data science backgrounds — came from Plista, an adtech startup, where they worked on products including an AI-powered ad creation tool. Haystack lets developers build pipelines for NLP use cases. Originally created for search applications, the framework can power engines that answer specific questions (e.g., “Why are startups moving to Berlin?”) or sift through documents. Haystack can also field “knowledge-based” searches that look for granular information on websites with a lot of data or internal wikis.
What strikes me? Three things:
- This is essentially a consulting and services approach
- Enterprise becomes apps for a situation, department, or specific need
- The buzzwords are interesting: NLP, semantic search, BERT, and humor.
Humor is a necessary quality which trying to make decades old technology work for distributed, heterogeneous data, email on a sales professionals mobile, videos, audio recordings, images, engineering diagrams along with the nifty datasets for the gizmos in the illustration, etc.
A question: Is $14 million enough?
Crickets.
Stephen E Arnold, April 29, 2022
Monopolies Know Best: The Amazon Method Involves a Better Status Page
December 13, 2021
Here’s the fix for the Amazon AWS outage: An updated status page. “Amazon Web Services Explains Outage and Will Make It Easier to Track Future Ones” reports:
A major Amazon Web Services outage on Tuesday started after network devices got overloaded, the company said on Friday [December 10, 2021] . Amazon ran into issues updating the public and taking support inquiries, and now will revamp those systems.
Several questions arise:
- How are those two pizza technical methods working out?
- What about automatic regional load balancing and redundancy?
- What is up with replicating the mainframe single point of failure in a cloudy world?
Neither the write up nor Amazon have answers. I have a thought, however. Monopolies see efficiency arising from:
- Streamlining by shifting human intermediated work to smart software which sort of works until it does not.
- Talking about technical prowess via marketing centric content and letting the engineering sort of muddle along until it eventually, if ever, catches up to the Mad Ave prose, PowerPoints, and rah rah speeches at bespoke conferences
- Cutting costs where one can; for example, robust network devices and infrastructure.
The AT&T approach is a goner, but it seems to be back, just in the form of Baby Bell thinking applied to an online bookstore which dabbles in national security systems and methods, selling third party products with mysterious origins, and promoting audio books to those who have cancelled the service due to endless email promotions.
Yep, outstanding, just from Wall Street’s point of view. From my vantage point, another sign of deep seated issues. What outfit is up next? Google, Microsoft, or some back office provider of which most humans have never heard?
The new and improved approach to an AT&T type business is just juicy with wonderfulness. Two pizzas. Yummy.
Stephen E Arnold, December 13, 2021
Semantics and the Web: A Snort of Pisco?
November 16, 2021
I read a transcript for the video called “Semantics and the Web: An Awkward History.” I have done a little work in the semantic space, including a stint as an advisor to a couple of outfits. I signed confidentiality agreements with the firms and even though both have entered the well-known Content Processing Cemetery, I won’t name these outfits. However, I thought of the ghosts of these companies as I worked my way through the transcript. I don’t think I will have nightmares, but my hunch is that investors in these failed outfits may have bad dreams. A couple may experience post traumatic stress. Hey, I am just suggesting people read the document, not go bonkers over its implications in our thumbtyping world.
I want to highlight a handful of gems I identified in the write up. If I get involved in another world-saving semantic project, I will want to have these in my treasure chest.
First, I noted this statement:
“Generic coding”, later known as markup, first emerged in the late 1960s, when William Tunnicliffe, Stanley Rice, and Norman Scharpf got the ideas going at the Graphics Communication Association, the GCA. Goldfarb’s implementations at IBM, with his colleagues Edward Mosher and Raymond Lorie, the G, M, and L, made him the point person for these conversations.
What’s not mentioned is that some in the US government became quite enthusiastic. Imagine the benefit of putting tags in text and providing electronic copies of documents. Much better than loose-leaf notebooks. I wish I have a penny for every time I heard this statement. How does the government produce documents today? The only technology not in wide use is hot metal type. It’s been — what? — a half century?
Second, I circled this passage:
SGML included a sample vocabulary, built on a model from the earliest days of GML. The American Association of Publishers and others used it regularly.
Indeed wonderful. The phrase “slicing and dicing” captured the essence of SGML. Why have human editors? Use SGML. Extract chunks. Presto! A new book. That worked really well but for one drawback: The proliferation of wild and crazy “books” were tough to sell. Experts in SGML were and remain a rare breed of cat. There were SGML ecosystems but adding smarts to content was and remains a work in progress. Yes, I am thinking of Snorkel too.
Third, I like this observation too:
Dumpsters are available in a variety of sizes and styles. To be honest, though, these have always been available. Demolition of old projects, waste, and disasters are common and frequent parts of computing.
The Web as well as social media are dumpsters. Let’s toss in TikTok type videos too. I think meta meta tags can burn in our cherry red garbage container. Why not?
What do these observations have to do with “semantics”?
- Move from SGML to XML. Much better. Allow XML to run some functions. Yes, great idea.
- Create a way to allow content objects to be anywhere. Just pull them together. Was this the precursor to micro services?
- One major consequence of tagging or the lack of it or just really lousy tagging, marking up, and relying of software allegedly doing the heavy lifting is an active demand for a way to “make sense” of content. The problem is that an increasing amount of content is non textual. Ooops.
What’s the fix? The semantic Web revivified? The use of pre-structured, by golly, correct mark up editors? A law that says students must learn how to mark up and tag? (Problem: Schools don’t teach math and logic anymore. Oh, well, there’s an online course for those who don’t understand consistency and rules.)
The write up makes clear there are numerous opportunities for innovation. And the non-textual information. Academics have some interesting ideas. Why not go SAILing or revisit the world of semantic search?
Stephen E Arnold, November 16, 2021
Facebook Targets Paginas Amarillas: Never Enough, Zuck?
October 14, 2021
Facebook is working to make one of its properties more profitable. The Next Web reports, “WhatsApp Reinvents the ‘Yellow Pages’ and Proves there Are No New Ideas.” The company will test out a new business directory feature in San Paulo, Brazil, where local users will be able to search for “businesses nearby” through the app. Writer Ivan Mehta reports:
“For years, Facebook and Instagram have been trying to connect you to businesses and make your shop through their platforms. While the WhatsApp Business app has been around, you couldn’t really search for businesses using the app, unless you’ve interacted with them previously. WhatsApp already offers payment services in Brazil. So it makes sense for it to provide discovery services for local businesses, so you can shop for goods in person, and pay through the platform. The chat app doesn’t have any ads, unlike Facebook and Instagram, so business interactions and transactions are one of the biggest ways for Facebook to earn some moolah out of it. In June, the company integrated its Shops feature in WhatsApp. So, we can expect more business-facing features in near future.”
India and Indonesia are likely next on the list for the project, according to Facebook’s Matt Idema. We are assured the company will track neither users’ locations nor the businesses they search for. Have we heard similar promises before?
Cynthia Murrell, October 14, 2021
Ex-Googlers Work On Biased NLP Solutions
October 6, 2021
Google is on top of the world when it comes to money and technology. Google is the world’s most used search engine, its Chrome Web browser is used by two-thirds of users, and about 29% of 2021 digital advertising were Google ads. Fast Company asks and investigates important questions about Google’s product quality in: “It’s Not Just You. Google Search Really Is Getting Worse.”
Over 80% of Alphabet Inc.’s revenue, Google’s parent company, comes from advertising revenue and about 85% of the world’s search engine traffic feeds through Google. Google controls a lot of users’ screen time. The search engine’s quality results have been studied and researchers have learned that very few users scroll past the “fold” (all of the available content on a screen). Advertising space at the top of search results is incredibly valuable. It also means that users are forced to scroll further and further to reach non-paid results.
Alphabet Inc. has another revenue generating platform, YouTube. A huge portion of videos include multiple ads. Users can avoid ads by paying for a premium subscription, but very few do.
Google does want to improve its search quality. Currently a lot of information from queries are distributed across multiple Web sites. Google wants to condense everything:
“Google is working on bringing this information together. The search engine now uses sophisticated “natural language processing” software called BERT, developed in 2018, that tries to identify the intention behind a search, rather than simply searching strings of text. AskJeeves tried something similar in 1997, but the technology is now more advanced.
BERT will soon be succeeded by MUM (Multitask Unified Model), which tries to go a step further and understand the context of a search and provide more refined answers. Google claims MUM may be 1,000 times more powerful than BERT, and be able to provide the kind of advice a human expert might for questions without a direct answer.”
Google controls a huge portion of the Internet and how users utilize it. Alphabet Inc. is here to stay for a long time, but there are alternatives such as Bing, DuckDuckGo, Ecosia, and Tor browsers. Google, however, will one day fade. Sears Roebuck, Blockbuster, Kmart, cassettes, etc. were al household names, until they became obsolete.
Whitney Grace, October 6, 2021

