Curated for content, computing, data, information, and digital experience professionals

Category: Semantic technologies (Page 26 of 72)

Our coverage of semantic technologies goes back to the early 90s when search engines focused on searching structured data in databases were looking to provide support for searching unstructured or semi-structured data. This early Gilbane Report, Document Query Languages – Why is it so Hard to Ask a Simple Question?, analyses the challenge back then.

Semantic technology is a broad topic that includes all natural language processing, as well as the semantic web, linked data processing, and knowledge graphs.


Convergence of Enterprise Search and Text Analytics is Not New

Prompted by the news item about IBM’s bid for SPSS and similar acquisitions by Oracle, SAP and Microsoft made me think about the predictions of more business intelligence (BI) capabilities being conjoined with enterprise search. But why now and what is new about pairing search and BI? They have always been complementary, not only for numeric applications but also for text analysis. Another article by John Harney in KMWorld referred to the “relatively new technology of text analytics” for analyzing unstructured text. The article is a good summary of some newer tools but the technology itself has had a long shelf life, too long for reasons which I’ll explore later.

Like other topics in this blog this one requires a readjustment in thinking by technology users. One of the great things about digitizing text was the promise of ways in which it could be parsed, sorted and analyzed. With heavy adoption of databases that specialized in textual, as well as numeric and date data fields for business applications in the 1960s and 70s, it became much easier for non-technical workers to look at all kinds of data in new ways. Early database applications leveraged their data stores using command languages; the better ones featured statistical analysis and publication quality report builders. Three that I was familiar with were DRS from ADM, Inc., BASIS from Battelle Columbus Labs and INQUIRE from IBM.

Tools that accompanied database back-ends had the ability to extract, slice and dice the database content, including very large text fields to report: word counts, phrase counts (breaking on any delimiter), transaction counts, relationships among data elements across associated record types, ability to create relationships on the fly, report expert activity and working documents, and describe distribution of resources. These are just a few examples of how new content assets could be created for export in minutes. In particular, a sort command with DRS had histogram controls that were invaluable to my clients managing corporate document and records collections, news clippings files, photographs, patents, etc. They could evaluate their collections by topic, date ranges, distribution, source, and so on, at any time.

So, there existed years ago the ability to connect data structures and use a command language to formulate new data models that informed and elucidated how information was being used in the organization, or to illustrate where there were holes in topics related to business initiatives. What were the barriers to wide-spread adoption? Upon reflection, I came to realize that extracting meaningful content from database in new and innovative formats requires a level of abstract thinking for which most employees are not well-trained. Putting descriptive data into a database via a screen form, then performing a transaction on the object of that data on another form, and then adding more data about another similar but different object are isolated in the database user’s experience and memory. The typical user is not trained to think about how the pieces of data might be connected in the database and therefore is not likely to form new ideas of how it can all be extracted in a report with new information about the content. There is a level of abstraction that eludes most workers whose jobs consist of a lot of compartmentalized tasks.

It was exciting to encounter prospects that really grasped the power of these tools and were excited to push the limits of the command language and reporting applications, but they were scarce. It turned out that our greatest use came in applying text analytics to the extraction of valuable information from our customer support database. A rigorously disciplined staff populated it after every support call with not only demographic information about the nature of the call, linked to a customer record that had been created back at the first contact during the sales process (with appropriate updates along the way in the procurement process) but also a textual description of the entire transaction. Over time this database was linked to a “wish list” database and another “fixes” database and the entire networked structure provided extremely valuable reports that guided both development work and documentation production. We also issued weekly summary reports to the entire staff so everyone was kept informed about product conditions and customer relationships. The reporting tools provided transparency to all staff about company activity and enabled an early version of “social search collaboration.”

Current text analytics products have significantly more algorithmic horsepower than the old command languages. But making the most of their potential and transforming them into utilities that any knowledge worker can leverage will remain a challenge for vendors in the face of poor abstract reasoning among much of the work force. The tools have improved but maybe not in all the ways they need to for widespread adoption. Workers should not have to be dependent on IT folks to create that unique analysis report that reveals a pattern or uncovers product flaws described by multiple customers. We expect workers to multitask, have many aptitudes and skills, and be self-servicing in so many aspects of their work, but for them to flourish the tools fall short too often. I’m putting in a big plug for text analytics for the masses, soon, so that enterprise search begins to deliver more than personalized lists of results for one person at a time. Give more reporting power to the user.

Semantic Search has Its Best Chance for Successes in the Enterprise

I am expecting significant growth in the semantic search market over the next five years with most of it focused on enterprise search. The reasons are pretty straightforward:

  • Semantic search is very hard and to scale it to the Web compounds the complexity.
  • Because the semantic Web is so elusive and results have been spotty with not much traction, it will be some time before it can be easily monetized.
  • Like many things that are highly complex, a good model will be to break the challenge of semantic search into smaller targeted business problems where focus is on a particular audience seeking content from a narrower domain.

I base this predication on my observation of the on-going struggle for organizations to get a strong framework in place to manage content effectively. By effectively I mean, establishing solid metadata, governance and publishing protocols that ensure that the best information knowledge workers produce is placed in range for indexing and retrieval. Sustained discipline and the people to exercise it just aren’t being employed in many enterprises to make this happen in a cohesive and comprehensive fashion. I have been discouraged by the number of well-intentioned projects I have seen flounder because organizations just can’t commit long-term or permanent human resources to the activity of content governance. Sometimes it is just on-again-off-again. What enterprises need are people with deep knowledge about the organization and how its content fits together in a logical framework for all types of knowledge workers. Instead, organizations tend to assign this job to external consultants or low-level staffers who are not well-grounded in the work of the particular enterprise. The results are predictably disappointing.

Enter semantic search technologies where there are multiple algorithmic tools available to index and retrieve content for complex and multi-faceted queries. Specialized semantic technologies are often well suited to shorter term projects for which domain specific vocabularies can be built more quickly with good results. Maintaining targeted vocabulary ontologies for a focused topic can be done with fewer human resources and a carefully bounded ontology can become an intelligent feed to a semantic search engine, helping it index with better precision and relevance.

This scenario is proposed with one caveat; enterprises must commit to having very smart people with enterprise expertise to build the ontology. Having a consultant coach the subject matter expert in method, process and maintenance guidelines for doing so is not a bad idea but the consultant has to prepare the enterprise for sustainability after exiting the scene.

The wager here is that enterprises can ramp up semantic search with a series of short, targeted projects, each of which establishes a goal of solving one business problem at a time and committing to efficient and accurate content retrieval as part of the solution. By learning what works well in each situation, intranet web retrieval will improve systematically and thoughtfully. The ramp to a better semantic Web will be paved with these interlocking pieces.

Keep an eye on these companies to provide technologies for point solutions in business critical applications: Basis Technology, Cognition Technology, Connotate, Expert Systems, Lexalytics, Linguamatics, Metatomix, Semantra, Sinequa and Temis.

Ontopia 5.0.0 Released

The first open source version of Ontopia has been released, which you can download from Google Code. This is the same product as the old commercial Ontopia Knowledge Suite, but with an open source license, and with the old license key restrictions removed. The new version has been created by not just by the Bouvet employees who have always worked on the product, but also by open source volunteers. In addition to bug fixes and minor changes, the main new features in this version are: Support for TMAPI 2.0; The new tolog optimizer; The new TologSpy tolog query profiler; The net.ontopia.topicmaps.utils. QNameRegistry and QNameLookup classes have been added, providing  lookup of topics using qnames; Ontopia now uses the Simple Logging Facade for Java (SLF4J), which makes it easier to switch logging engines, if desired. http://www.ontopia.net/

Lucid Imagination and ISYS Partner on Lucene/Solr

Lucid Imagination and ISYS Search Software announced a strategic partnership. The agreement enables Lucid Imagination to provide solutions that combine its core Lucene and Solr expertise with the ISYS File Readers document filtering technology. The flexibility of the architecture allows enterprises to develop sophisticated purpose-built search solutions. By offering ISYS File Readers as part of its Lucene/Solr solutions, Lucid Imagination gives users and developers out-of-the-box capability to find and extract virtually all of the content and formats that exist in their enterprise environment. Lucid Imagination Web site serves as a knowledge portal for the Lucene community, with wide range of information, resources and  information retrieval application, LucidFind to help developers and search professionals get access to the information they need to design, build and deploy Lucene and Solr based solutions. http://www.lucidimagination.com

MuseGlobal and Specialty Systems Partner

MuseGlobal announced a partnership with Specialty Systems, Inc., a company focusing on innovative information systems solutions to Federal, State and Local Government customers. Specialty Systems, Inc. is partnering with MuseGlobal to provide the systems integration expertise to engineer law enforcement and homeland security applications built on MuseGlobal’s MuseConnect, which provides federated search and harvesting technologies, with a library of more than 6,000 pre-built source connectors. The applications resulting from this partnership will incorporate unified information access allowing structured data from database sources; semi-structured data from spreadsheets, forms and XML sources; unstructured data from web sites, documents, email; and rich media such as images, video and audio information to be accessed simultaneously from internal databases and external sources.  This information is gathered on the fly, and unified for immediate presentation to the requestor. http://www.specialtysystems.com, http://www.museglobal.com

SDL Tridion Integrates Q-go Natural Language Search into Web Content Management

SDL Tridion announced that it has partnered with Q-go to provide an integrated Natural Language Search engine within SDL Tridion’s web content management platform. The solution provides the online search environment within websites only targeted and relevant search results. Q-go’s Natural Language Search is now accessible from within the SDL Tridion web content management environment. Content editors are able to create model questions in the Q-go component of the SDL Tridion platform. This means that the most common questions pertaining to products and the website itself can be targeted and answered by web content editors, creating streamlined content and vastly increased relevance of searches. The integration also means that only one interface is needed to update the entire website, which can be done anywhere, anytime. You can find more information on the integration at the eXtensions Community of  http://www.sdltridionworld.com

If a Vendor Spends Enough on Full-page Ads: Ink will Follow

Earlier comments in this blog referred to Autonomy ads in Information Week. They have continued throughout early 2009 with just the latest proclaiming “Autonomy Dominates Enterprise Search” in bold red and black, two of my favorite, eye-catching colors. Having read the publication for over ten years, I notice things that are different. Seeing a search company repeatedly showing up keeps me noticing because they are the first to spend on major advertising like this in an IT publication.

This week the predictable happened, it was an article by Information Week‘s Sr. VP focusing on Autonomy’s terrific business run in a tough economy. Fair enough – it happens all the time for big spenders.

I just want to remind readers that if you are a small unit in a large organization or a small or medium business, there are dozens of enterprise search solutions that will serve you extremely well, with much lower cost of ownership and startup effort than Autonomy. You do not need the biggest or fastest growing company’s products to get good or even excellent solutions. Furthermore, the chances of getting superior customer support and services from a more modest company, which is focused exclusively on search excellence, are much better.

Be sure to check out the offerings at the Gilbane Conference in San Francisco next week. A lot more guidance and good case studies will give you an earful of what else to consider. The search headliners at the conference with Hadley Reynolds moderating are:

E8. Search Survival Guide: Delivering Great Results
Speakers: Randy Woods, Co-founder & Executive VP, non-linear creations, Best Practices for Tuning Enterprise Search and Miles Kehoe, President, New Idea Engineering

E9/I5. The Next Big Thing: Tomorrow’s Search Revealed
Speakers: Stephen Arnold, ArnoldIT, What You Need to Know About Google Dataspaces and Jeff Fried, Senior Product Manager, Microsoft

E10/I6. Bringing it All Together: Perils and Pitfalls of Search Federation
Speakers: Helen Mitchell Curtis, Senior Program Director of Enterprise Solutions, MacFadden, Federated Search in a Disparate Environment, Larry Donahue, Chief Operating Officer & Corporate Counsel, Deep Web Technologies, Federated Search: True Enterprise Search and Jeff Fried, Senior Product Manager, Microsoft

E11/I7. The Special Case of Categories – and Where To Find Them
Speakers: Joseph Busch, Founder, Taxonomy Strategies, Taxonomy Validation, and Arje Cahn, CTO, Hippo, Find What You Need in Unstructured Content with the Help of Others (and your CMS): Demo of Wikipedia with Faceted Search

E12/I8. It’s Easier with Structure: Leveraging Markup for Better Search
Speakers: Dianne Burley, Industry Specialist, Nstein Technologies, Semantic Search and J. Brooke Aker, CEO, Expert System, A 3-Step Walk Through ECM Using Semantics

E13/I9. Improving SharePoint Search & Navigation with a Taxonomy and Metadata

Have a question for our analyst panel?

Looking forward to seeing many of you next week at Gilbane San Francisco. Whether you will be there or not, you can suggest questions to ask our analyst panel. Each of the panelists have specific areas of expertise covering web content management, web governance, enterprise social software and social media, collaboration, and enterprise search. The panel is a keynote session after the two keynote presentations from Microsoft and Adobe, so we’ll also be covering reactions to those. You can submit your questions directly to me via a comment, email, or twitter (DM or post using the hashtag #gilbanesf).

Registration for the conference is still open and will be available on-site. If you register in advance you can still get a $200. discount using GILBANE as the discount code. There is no charge for the keynotes or the technology demonstrations or product labs.

K2. Keynote Analyst Panel
We invite industry analysts from many different firms to speak at all our events to make sure our conference attendees hear differing opinions from a wide variety of expert sources. A second, third, fourth or fifth opinion will ensure you don’t make ill-informed decisions about critical content and information technologies or strategies. This session will be a lively, interactive debate guaranteed to be both informative and fun.
Moderator: Frank Gilbane
Panelists:
Jeremiah Owyang, Senior Analyst, Social Computing, Forrester
Hadley Reynolds, Research Director, Search & Digital Marketplace Technologies, IDC
Larry Hawes, Lead Analyst, Collaboration and Enterprise Social Software, Gilbane Group
Lisa Welchman, Founding Partner, WelchmanPierpoint

« Older posts Newer posts »

© 2026 The Gilbane Advisor

Theme by Anders NorenUp ↑