Who Blew the Whistle on BlackCore
June 23, 2026
Another dinobaby post. No AI unless it is an image. This dinobaby is not Grandma Moses, just Grandpa Arnold.
One of the more interesting Israeli intelware service firms is BlackCore. I know most people have not heard of the company and the firm does not exactly put up billboards on Route 101 from San Francisco to the Plastic Fantastic of Silicon Valley. However, information from France about a firm that does the Cambridge Analytica social media manipulation thing has been reported in stories from Russia to Yahoo News (“Secrets of BlackCore: Israel’s Information Warfare Now Has Few Limits or Boundaries”). In the US the news appears to pivot on sports, and that is one interesting characteristic of knowledge flow in the US of A today.

Thanks, Venice.ai. You did not hang, time out, or tell me my image prompt violated your perception of proper generated imagery. Quite a surprise!
BlackCore, also rumored to do business as Mycelium, uses a range of methods to shape or weaponized certain information. The approach is standard operating procedure for lots of companies, independent contractors, and PR outfits. Therefore, it is difficult for me to get excited about this flurry of news stories. BlackCore and its inauthentic information angle may have blipped French radar during recent elections. But I am in rural Kentucky and relying on my experience about what catches the eye of French intelligence professionals. In general, it is not a great idea to light up their radar. Things happen. Case in point: The flood of articles outing the hard working team at BlackCore and its affiliates.
I would point out that some of the information may have originated intentionally or by “accident” from an intel outfit in the French government. One of my team told me that the original chunk of information about BlackCore evolved from a unit possibly affiliated with Secretariat-General for Defense and National Security (SGDSN). This outfit reports directly to the French prime minister.
Several questions occur to me:
- Why is Meta fixated on NSO Group? That is, in my opinion, a precursor to more sophisticated intelware operations. Perhaps Meta does not know who is doing interesting things on Meta. The Zuck just falls back on his perception about the Pegasus crowd.
- What sparked the flurry of news stories about an outfit that has done a pretty good job of remaining off the grid occupied by journalists, do-gooders, and civilians? Why now?
- What will be done about BlackCore and similar entities operating in other countries; for example, China and Russia? My thought is that not much has been done and not much will be done. The fact that Cambridge Analytica left a game plan for every misinformation operation worldwide is one of its stellar contributions. Why was it allowed to operate with considerable freedom for several years?
Net net: Snappy name aside, BlackCore is just one of a flock of similar intelware operations. A bit of poking around on LinkedIn-type services provide plenty of hints about who is doing what.
PS. Last time I was in France I tried to get a hat with the entity name “Viginum” on it. I failed.
Stephen E Arnold, June 23, 2026
Harper: Price Is a Differentiator. Right, Grammarly?
June 18, 2026
Harper is a free grammar checker that rivals Grammarly.
Grammar, punctuation, and all those editing skills are still necessary and important to the English language. Without proper rules, then there would not be a universal understanding of what it means to speak and read English. Those are grandiose thoughts, but even with today’s spell checkers and grammar editors there is still room for improvement.? ?
This is true when comes to SaaS tool like Grammarly. These tools are amazing, but paying for what was once free is a big bummer. Thankfully a developer uploaded Harper to GitHub.
“Harper is a free, open-source grammar checker designed to be just right. Think of it as the private alternative to Grammarly, built after years of dealing with the shortcomings of the competition. Harper catches the kinds of mistakes that matter: improper capitalization, misspelled words, awkward phrasing, and broken grammar. Your writing never leaves your computer.”
I forgot to mention that Grammarly and other subscription grammar checkers, remember and send out your information. It’s spying on your writer.? ? Harper is an intuitive grammar and spell checker that is on par with anything that requires payment. Plus it doesn’t spy on you.
It catches simple mistakes and makes suggestions on how to improve writing. It checks tone, counts spaces between words, and understands sentence structure and meaning. It’s not 100% foolproof, because it does misunderstand the usage of some words and their placement in sentence, but Harper only flags those for review. Harper runs locally and is downloadable for multiple OSs and platforms. It’s a great grammar editor to check out and I’d recommend it, especially if you do a lot of writing and your computer lacks a robust grammar engine. I’m looking at you Apple.
Whitney Grace, June 18, 2026
No Joke: BS Is a Signal for Work Competence
April 1, 2026
How many of us have worked at a job and met an individual who used more jargon than a dictionary? Everyone! Do you remember that suspicion at the back of your thoughts that this person was crappy at their tasks? BINGO! Cornell University developed the “the Corporate Bullshit Receptivity Scale, a tool designed to measure how impressed people are by business school-style jargon that sounds strategic but says very little.” The Register discusses it in the article, “Those Who ‘Circle Back’ And ‘Synergize’ Also Tend To Be Crap At Their Jobs.”
Cornell studied corporate jargon and discovered people who found it to be helpful are more likely to struggle with workplace decision-making and analytical thinking. Here’s the just of Cornell’s research:
“To build the scale, researchers ran four studies involving more than 1,000 working adults in the US and Canada. Participants were shown a mix of genuine corporate statements and nonsense lines generated by what the researchers call a “corporate bullshit generator” – effectively a tool that mashes together buzzwords into sentences that sound like they came straight out of a quarterly strategy meeting.”
The participants were asked to rate how useful the statements were or how someone interprets impressive-sounding language. The results weren’t flattering and corresponded to some less than basic cognitive patterns.
Here’s the results:
“People who scored higher on the Corporate Bullshit Receptivity Scale tended to perform worse on tests measuring analytical thinking, cognitive reflection, and fluid intelligence. They also made poorer judgments in workplace decision-making scenarios designed to mimic common business problems. In other words, the employees most impressed by corporate jargon were also the ones least likely to think critically about it.”
Those who found this buzzword users as charismatic and inspiring also used the jargon themselves. Let’s assume that this research nails the obvious as insecurity and the imposter syndrome vie for dominance. [Please, insert your own jargon-filled statement here. Thank you.]
Whitney Grace, April 1, 2026
Big Tech AI: Biased or Not?
March 20, 2026
Another dinobaby post. No AI unless it is an image. This dinobaby is not Grandma Moses, just Grandpa Arnold.
I read an unusual news item published by Versant CNBC. Its title is “Anthropic’s Claude Would Pollute Defense Supply Chain: Pentagon CTO.” I don’t know much about the US government and I know even less about the Department of War. What I do know is that Versant CNBC called attention to a facet of smart software most ignored. Dr. Timnit Gebru raised some questions about AI bias, and she was invited to find her future elsewhere along with her pet stochastic parrot. Others have suggested that certain content is under-represented; namely, I have with regard to coverage of information in other major countries. To get Chinese and Russian perspectives, I have to use language specific indexes and rely on online translation services. The information I have located is not well represented in result sets my team and I have reviewed from US big tech outfits’ AI systems. Yeah, English and low-hanging fruit are more common, not a salient post from a Chinese or Russian language source. Your mileage may vary, but I am a dinobaby, and I don’t wander too far from the outmoded ideas about editorial policies, precision, recall, and other other impedimenta from ancient online services.

A smart software system is testifying about policy biases before a distinguished body of elected officials. Thanks, Venice.ai. Good enough.
The Versant CNBC outfit which I will refer to as VC NBC reports:
Defense Department CTO Emil Michael on Thursday [March 12, 2026] said Anthropic’s Claude artificial intelligence models would “pollute” the agency’s supply chain because they have “a different policy preference” that is baked in.
Okay, “a different policy preference” suggests to me:
- The developers of smart software can steer what the models output; that is, weaponize them, shape them, make them formulate responses that affect the systems or users ingesting AI output
- Professionals at in the US government have determined from their own observations and by consulting trusted experts like those from Palantir Technologies that their determinations are accurate and valid based on the systems in use prior to this determination
- Users of these systems and analysts of these systems have not been sufficiently critical of AI outputs to observe these pollutive functions and the explicit policy preferences noted in the information presented by VC NBC.
Let’s assume the information presented by VC NBC is spot on. I have several questions:
- Is it easy to shape or weaponize the probabilistic word guessing systems to make a duck the equivalent of a cow or to present a fact as an incorrect assertion? If yes, who are the experts turning the knobs and twisting the dials in these AI companies?
- Is one company capable of weaponizing and shaping, or can other AI outfits perform a similar calibration mechanism? If yes, are the Chinese and French models weaponized, shaped, or directed in a similar way? Other than academics publishing in ArXhiv, are there mainstream research outfits tracking these clever meta-editorial activities? Are RAND- or McKinsey-type outfits chasing this concept.
- Short of shutting down an alleged weaponizer, how will this “policy” shaping system be controlled? I think that may be difficult because some organizations like Microsoft have integrated an alleged policy shaper along with the allegedly more objective services. Note that the smart software, including the Chinese and French systems, do not include Chinese or Russian content. English seems to be the go-to language for training data. That decision may inject cultural bias I would suggest.
Net net: I think that policy shaping is now a “fact” that may have some persistence in the AI world. I will be interested in watching how the AI firms explain and demonstrate that the outputs of their systems are not just “correct” but “objective” according to the standard used by the US government. Does anyone care? That is an important question. I want to avoid Zitron-onics and say, “Worth monitoring.”
Stephen E Arnold, March 20, 2026
FAIR Squared Data Management Promises to Find Missing Data
October 31, 2025
It’s very true that information is lost or hidden away in archives never to see the light of day. That’s why it’s important to preserve the information and even use AI to make it available. Science Daily reports on new information management tool that claims to have a solution: “90% Of Science Is Lost. This New AI Just Found It.” FAIR² Data Management is designed by Frontiers and is:
“…described as the world’s first comprehensive, AI-powered research data service. It is designed to make data both reusable and properly credited by combining all essential steps — curation, compliance checks, AI-ready formatting, peer review, an interactive portal, certification, and permanent hosting — into one seamless process. The goal is to ensure that today’s research investments translate into faster advances in health, sustainability, and technology.”
The data management system is built on a robust AI algorithm. Researchers feed their their data into FAIR² and four integrated outputs are returned: a certificate, an interactive data portal with AI chat and visualizations, peer-reviewed and citable data article, and a certified data package. All of these components work “[t]ogether,…to…ensure that every dataset is preserved, validated, citable, and reusable, helping accelerate discovery while giving researchers proper recognition.”
This is a great idea and how AI should ideally be used to ensure that information is credible. If only all AI algorithms employed a data management algorithm like this to prevent AI slop, drivel, and garbage from clogging up the Internet and our brains.
Whitney Grace, October 31, 2025
Text Wranglers, Attention
October 13, 2025
This is a short item for people who manipulate or wrangle text. Navigate to TextTools. The site provides access to several dozen utilities. I checked a handful of the services and found them to be free. The one I tested was Difference Checker. Paste the text of the two files or in my case code snippets. The output flags the differences. Worth a look.
Stephen E Arnold, October 13, 2025
Weaponization of LLMs Is a Thing. Will Users Care? Nope
October 10, 2025
This essay is the work of a dumb dinobaby. No smart software required.
A European country’s intelligence agency learned about my research into automatic indexing. We did a series of lectures to a group of officers. Our research method, the results, and some examples preceded a hands on activity. Everyone was polite. I delivered versions of the lecture to some public audiences. At one event, I did a live demo with a couple of people in the audience. Each followed a procedure, and I showed the speed with which the method turned up in the Google index. These presentations took place in the early 2000s. I assumed that the behavior we discovered would be disseminated and then it would diffuse. It was obvious that:
- Weaponized content would be “noted” by daemons looking for new and changed information
- The systems were sensitive to what I called “pulses” of data. We showed how widely used algorithms react to sequences of content
- The systems would alter what they would output based on these “augmented content objects.”
In short, online systems could be manipulated or weaponized with specific actions. Most of these actions could be orchestrated and tuned to have maximum impact. One example in my talks was taking a particular word string and making it turn up in queries where one would not expect that behavior. Our research showed that a few as four weaponized content objects orchestrated in a specific time interval would do the trick. Yep, four. How many weaponized write ups can my local installation of LLMs produce in 15 minutes? Answer: Hundreds. How long does it take to push those content objects into information streams used for “training.” Seconds.
Fish live in an environment. Do fish know about the outside world? Thanks, Midjourney. Not a ringer but close enough in horseshoes.
I was surprised when I read “A Small Number of Samples Can Poison LLMs of Any Size.” You can read the paper and work through the prose. The basic idea is that selecting or shaping training data or new inputs to recalibrate training data can alter what the target system does. I quite like the phrase “weaponize information.” Not only does the method work, it can be automated.
What’s this mean?
The intentional selection of information or the use of a sample of information from a domain can generate biases in what the smart software knows, thinks, decides, and outputs. Dr. Timnit Gebru and her parrot colleagues were nibbling around the Google cafeteria. Their research caused the Google to put up a barrier to this line of thinking. My hunch is that she and her fellow travelers found that content that is representative will reflect the biases of the authors. This means that careful selection of content for training or updating training sets can be steered. That’s what the Anthropic write up make clear.
Several observations are warranted:
- Whoever selects training data or the information used to update and recalibrate training data can control what is displayed, recommended, or included in outputs like recommendations
- Users of online systems and smart software are like fish in a fish bowl. The LLM and smart software crowd are the people who fill the bowl and feed the fish. Fish have a tough time understanding what’s outside their bowl. I don’t like the word “bubble” because these pop. An information fish bowl is tough to escape and break.
- As smart software companies converge into essentially an oligopoly using the types of systems I described in the early 2000s with some added sizzle from the Transformer thinking, a new type of information industrial complex is being assembled on a very large scale. There’s a reason why Sam AI-Man can maintain his enthusiasm for ChatGPT. He sees the potential of seemingly innocuous functions like apps within ChatGPT.
There are some interesting knock on effects from this intentional or inadvertent weaponization of online systems. One is that the escalating violent incidents are an output of these online systems. Inject some René Girard-type content into training data sets. Watch what those systems output. “Real” journalists are explaining how they use smart software for background research. Student uses online systems without checking to see if the outputs line up with what other experts say. What about investment firms allowing smart software to make certain financial decisions.
Weaponize what the fish live in and consume. The fish are controlled and shaped by weaponized information. How long has this quirk of online been known? A couple of decades, maybe more. Why hasn’t “anything” been done to address this problem? Fish just ask, “What problem?”
Stephen E Arnold, October x, 2025
I spotted
Content Injection Can Have Unanticipated Consequences
February 24, 2025
The work of a real, live dinobaby. Sorry, no smart software involved. Whuff, whuff. That’s the sound of my swishing dino tail. Whuff.
Years ago I gave a lecture to a group of Swedish government specialists affiliated with the Forestry Unit. My topic was the procedure for causing certain common algorithms used for text processing to increase the noise in their procedures. The idea was to input certain types of text and numeric data in a specific way. (No, I will not disclose the methods in this free blog post, but if you have a certain profile, perhaps something can be arranged by writing benkent2020 at yahoo dot com. If not, well, that’s life.)
We focused on a handful of methods widely used in what now is called “artificial intelligence.” Keep in mind that most of the procedures are not new. There are some flips and fancy dancing introduced by individual teams, but the math is not invented by TikTok teens.
In my lecture, the forestry professionals wondered if these methods could be used to achieve specific objectives or “ends”. The answer was and remains, “Yes.” The idea is simple. Once methods are put in place, the algorithms chug along, some are brute force and others are probabilistic. Either way, content and data injections can be shaped, just like the gizmos required to make kinetic events occur.
The point of this forestry excursion is to make clear that a group of people, operating in a loosely coordinated manner can create data or content. Those data or content can be weaponized. When ingested by or injected into a content processing flow, the outputs of the larger system can be fiddled: More emphasis here, a little less accuracy there, and an erosion of whatever “accuracy” calculations are used to keep the system within the engineers’ and designers’ parameters. A plebian way to describe the goal: Disinformation or accuracy erosion.
I read “Meet the Journalists Training AI Models for Meta and OpenAI.” The write up explains that journalists without jobs or in search of extra income are creating “content” for smart software companies. The idea is that if one just does the Silicon Valley thing and sucks down any and all content, lawyers might come calling. Therefore, paying for “real” information is a better path.
Please, read the original article to get a sense of who is doing the writing, what baggage or mind set these people might bring to their work.
If the content is distorted — either intentionally or unintentionally — the impact of these content objects on the larger smart software system might have some interesting consequences. I just wanted to point out that weaponized information can have an impact. Those running smart software and buying content assuming it is just fine, might find some interesting consequences in the outputs.
Stephen E Arnold, February 24, 2025
"Real" Entities or Sock Puppets? A New Solution Can Help Analysts and Investigators
January 28, 2025
Bitext’s NAMER (shorthand for "named entity recognition") can deliver precise entity tagging across dozens of languages.
Graphs — knowledge graphs and social graphs — have moved into the mainstream since Leonhard Euler formed the foundation for graph theory in the mid 18th century in Berlin.
With graphs, analysts can take advantage of smart software’s ability to make sense of Named Entity Recognition (NER), event extraction, and relationship mapping.
The problem is that humans change their names (handles, monikers, or aliases) for many reasons: Public embarrassment, a criminal record, a change in marital status, etc.
Bitext’s NER solution, NAMER, is specifically designed to meet the evolving needs of knowledge graph companies, offering exceptional features that tackle industry challenges.
Consider a person disgraced with involvement in a scheme to defraud investors in an artificial intelligence start up. The US Department of Justice published the name of a key actor in this scheme. (Source: https://www.justice.gov/usao-ndca/pr/founder-and-former-ceo-san-francisco-technology-company-and-attorney-indicted-years). The individual was identified by the court as Valerie Lau Beckman. The official court documents used the name "Lau" to reference her involvement in a multi-million dollar scam.
However, in order to correctly identify her in social media, subsequent news stories, and in possible public summaries of her training on a LinkedIn-type of smart software is not enough.
That’s the role of a specialized software solution. Here’s what NAMER delivers.
The system identifies and classifies entities (e.g., people, organizations, locations) in unstructured data. The system accurately links data across different sources of content. The NAMER technology can tag and link significant events (transactions, announcements) to maintain temporal relevance; for example, when Ms. Lau Beckman is discharged from the criminal process. NAMER can connect entities like Ms. Lau or Ms. Beckman to other individuals with whom she works or interacts and her "names" appearance in content streams.
The licensee specifies the languages NAMER is to process, either in a knowledge base or prior to content processing via a large language model.
Access to the proprietary NAMER technology is via a local SDK which is essential for certain types of entity analysis. NAMER can also be integrated into another system or provided as a "white label service" to enhance an intelligence system with NAMER’s unique functions. The developer provides for certain use cases direct access to the source code of the system.
For an organization or investigative team interested in keeping data about Lau Beckman at the highest level of precision, Bitext’s NAMER is an essential service.
Stephen E Arnold, January 28, 2025
More about NAMER, the Bitext Smart Entity Technology
January 14, 2025
A dinobaby product! We used some smart software to fix up the grammar. The system mostly worked. Surprised? We were.
We spotted more information about the Madrid, Spain based Bitext technology firm. The company posted “Integrating Bitext NAMER with LLMs” in late December 2024. At about the same time, government authorities arrested a person known as “Broken Tooth.” In 2021, an alert for this individual was posted. His “real” name is Wan Kuok-koi, and he has been in an out of trouble for a number of years. He is alleged to be part of a criminal organization and active in a number of illegal behaviors; for example, money laundering and human trafficking. The online service Irrawady reported that Broken Tooth is “the face of Chinese investment in Myanmar.”
Broken Tooth (né Wan Kuok-koi, born in Macau) is one example of the importance of identifying entity names and relating them to individuals and the organizations with which they are affiliated. A failure to identify entities correctly can mean the difference between resolving an alleged criminal activity and a get-out-of-jail-free card. This is the specific problem that Bitext’s NAMER system addresses. Bitext says that large language models are designed for for text generation, not entity classification. Furthermore, LLMs pose some cost and computational demands which can pose problems to some organizations working within tight budget constraints. Plus, processing certain data in a cloud increases privacy and security risks.
Bitext’s solution provides an alternative way to achieve fine-grained entity identification, extraction, and tagging. Bitext’s solution combines classical natural language processing solutions solutions with large language models. Classical NLP tools, often deployable locally, complement LLMs to enhance NER performance.
NAMER excels at:
- Identifying generic names and classifying them as people, places, or organizations.
- Resolving aliases and pseudonyms.
- Differentiating similar names tied to unrelated entities.
Bitext supports over 20 languages, with additional options available on request. How does the hybrid approach function? There are two effective integration methods for Bitext NAMER with LLMs like GPT or Llama are. The first is pre-processing input. This means that entities are annotated before passing the text to the LLM, ideal for connecting entities to knowledge graphs in large systems. The second is to configure the LLM to call NAMER dynamically.
The output of the Bitext system can generate tagged entity lists and metadata for content libraries or dictionary applications. The NAMER output can integrate directly into existing controlled vocabularies, indexes, or knowledge graphs. Also, NAMER makes it possible to maintain separate files of entities for on-demand access by analysts, investigators, or other text analytics software.
By grouping name variants, Bitext NAMER streamlines search queries, enhancing document retrieval and linking entities to knowledge graphs. This creates a tailored “semantic layer” that enriches organizational systems with precision and efficiency.
For more information about the unique NAMER system, contact Bitext via the firm’s Web site at www.bitext.com.
Stephen E Arnold, January 14, 2025

