Category: Semantic technologies (Page 34 of 72)

Our coverage of semantic technologies goes back to the early 90s when search engines focused on searching structured data in databases were looking to provide support for searching unstructured or semi-structured data. This early Gilbane Report, Document Query Languages – Why is it so Hard to Ask a Simple Question?, analyses the challenge back then.

Semantic technology is a broad topic that includes all natural language processing, as well as the semantic web, linked data processing, and knowledge graphs.

Taxonomy and Enterprise Search

February 22, 2008 / Lynda Moulton / 1 Comment

This blog entry on the “Taxonomy Watch” website prompts me to correct the impression that I believe naysayers who say that taxonomies take too much time and effort to be valuable. Nothing could be further from the truth. I believe in and have always been highly vested in taxonomies because I am convinced that an investment in pre-processing enterprise generated content into meaningfully organized results brings large returns in time savings for a searcher. S/he, otherwise, needs to invest personally in the laborious post-processing activity of sifting and rejecting piles of non-relevant content. Consider that categorizing content well and only once brings benefit repeatedly to all who search an enterprise corpus.

Prime assets of enterprises are people and their knowledge; the resulting captured information can be leveraged as knowledge assets (KA). However, there is a serious problem “herding” KA into a form that results in leveragable knowledge. Bringing content into a focus that is meaningful to a diverse but specialized audience of users, even within a limited company domain is tough because the language of the content is so messy.

So, what does this have to do with taxonomies and enterprise search, and how they factor into leveraging KA? Taxonomies have a role as a device to promote and secure the meaningful retrievability of content when we need it most or fastest, just-in-time retrieval. If no taxonomies exist to pre-collocate and contextualize content for an audience, we will be perpetually stuck in a mode of having to do individual human filtering of excessive search results that come from “keyword” queries. If we don’t begin with taxonomies for helping search engines categorize content, we will certainly never get to the holy grail of semantic search. We need every device we can create and sustain to make information more findable and understandable; we just don’t have time to both filter and read, comprehensively, everything a keyword search throws our way to gain the knowledge we need to do our jobs.

Experts recognize that organizing content with pre-defined terminology (aka controlled vocabularies) that can be easily displayed in an expandable taxonomic structure is a useful aid for a certain type of searcher. The audience for navigated search is one that appreciates the clustering of search results into groups that are easily understood. They find value in being able to move easily from broad concepts to narrower ones. They especially like it when the categories and terminology are a close match to the way they view a domain of content in which they are subject experts. It shows respect for their subject area and gives them a level of trust that those maintaining the repository know what they need.

Taxonomies, when properly employed, serve triple duty. Exposing them to search engines that are capable of categorizing content puts them into play as training data. Setting them up within content management systems provides a control mechanism and validation table for human assigned metadata. Finally, when used in a navigated search environment, they provide a visual map of the content landscape.

U.S. businesses are woefully behind in “getting it;” they need to invest in search and surrounding infrastructure that supports search. Comments from a recent meeting I attended reflected the belief that the rest of the world is far ahead in this respect. As if to highlight this fact, a colleague just forwarded this news item yesterday. “On February 13, 2008, the XBRL-based financial listed company taxonomy formulated by the Shanghai Stock Exchange (SSE) was “Acknowledged” by the XBRL International. The acknowledgment information has been released on the official website of the XBRL International (http://www.xbrl.org/FRTaxonomies/)….”.

So, let’s get on with selling the basic business case for taxonomies in the enterprise to insure that the best of our knowledge assets will be truly findable when we need them.

Search Engines Under the Hood

February 12, 2008 / Lynda Moulton / 1 Comment

This week’s thoughts come from the pile of serendipitous reading that routinely piles up on my desk. In this case a short article in Information Week caught my eye because it featured the husband of a former neighbor, Ken Krugler, co-founder of Krugle. I’d set it aside because a fellow, David Eddy, in my knowledge management forum group keeps telling us that we need tools to facilitate searching for old but still useful source code. In order to do it, he believes, we need an investment in semantic search tools that normalize the voluminous language variants scattered throughout source code. That would enable programmers to find code that could be re-purposed in new applications.

Now, I have taken the position that source code is just one set of intellectual property (IP) asset that is wasted, abandoned and warehoused for technology archaeologists of centuries hence. I just don’t see a solid business case being made to develop search tools that will become a semantic search engine for proprietary treasure troves of code.

Enters old acquaintance Ken Krugler with what seems to be, at first glance, a Web search system that might be helpful for finding useful code out on the Web, including open source. I have finally visited his Web site and I see language and new offerings that intrigue me. “Krugle Enterprise is a valuable tool for anyone involved in software development. Krugle makes software development assets easily accessible and increases the value of a company’s code base. By providing a normalized view into these assets, wherever they may be stored, Krugle delivers value to stakeholders throughout the enterprise.” They could be onto something big. This is a kind of enterprise search I haven’t really had time to think about but may-be I will now.
One thing leading to another, I checked out Ken Krugler’s blog and saw an earlier posting: Is Writing Your Own Search Engine Hard? This is recommended reading for anyone who even dabbles in enterprise search technology but doesn’t want to get her/his hands dirty with the mechanics. It is short, to-the-point and summarizes how and why so many variations of search are battling it out in the marketplace.

I don’t want end-users to struggle too much with the under the hood details but when you are thinking about enterprise search for your organization, it is worth considering how much technology you are getting for the value you want it to deliver, year after year, as your mountains of IP content accrue. Don’t give this idea short shrift because search is an investment that keeps giving if it is chosen appropriately for the problem you need to solve.

Search Behind the Firewall aka Enterprise Search

February 1, 2008 / Lynda Moulton / 0 Comments

Called to account for the nomenclature “enterprise search,” which is my area of practice for The Gilbane Group, I will confess that the term has become as tiresome as any other category to which the marketplace gives full attention. But what is in a name, anyway? It is just a label and should not be expected to fully express every attribute it embodies. A year ago I defined it to mean any search done within the enterprise with a primary focus of internal content. “Enterprise” can be an entire organization, division, or group with a corpus of content it wants to have searched comprehensively with a single search engine.

A search engine does not need to be exclusive of all other search engines, nor must it be deployed to crawl and index every single repository in its path to be referred to as enterprise search. There are good and justifiable reasons to leave select repositories un-indexed that go beyond even security concerns, implied by the label “search behind the firewall.” I happen to believe that you can deploy enterprise search for enterprises that are quite open with their content and do not keep it behind a firewall (e.g. government agencies, or not-for-profits). You may also have enterprise search deployed with a set of content for the public you serve and for the internal audience. If the content being searched is substantively authored by the members of the organization or procured for their internal use, enterprise search engines are the appropriate class of products to consider. As you will learn from my forthcoming study, Enterprise Search Markets and Applications: Capitalizing on Emerging Demand, and that of Steve Arnold (Beyond Search) there are more than a lot of flavors out there, so you’ll need to move down the food chain of options to get it right for the application or problem you are trying to solve.

OK! Are you yet convinced that Microsoft is pitting itself squarely against Google? The Yahoo announcement of an offer to purchase for something north of $44 billion makes the previous acquisition of FAST for $1.2 billion pale. But I want to know how this squares with IBM, which has a partnership with Yahoo in the Yahoo edition of IBM’s OmniFind. This keeps the attorneys busy. Or may-be Microsoft will buy IBM, too.

Finally, this dog fight exposed in the Washington Post caught my eye, or did one of the dogs walk away with his tail between his legs? Google slams Autonomy – now, why would they do that?

I had other plans for this week’s blog but all the Patriots Super Bowl talk puts me in the mode for looking at other competitions. It is kind of fun.

Beyond Search and Search

February 1, 2008 / Frank Gilbane / 0 Comments

As many of you know from our press release at Gilbane Boston, two of the reports we will be publishing in the next few of months have to do with search. Lynda Moulton, who runs our Enterprise Search consulting practice is working on Enterprise Search Markets and Applications: Capitalizing on Emerging Demand, and our colleague Steve Arnold is writing Beyond Search: What to do When you’re Enterprise Search System Doesn’t Work. Lynda’s report covers the “Enterprise Search” market, what organizations are doing with the variety of technologies considered to be enterprise search products, and what their experiences have been. By the way Lynda is collecting experiences about implementations and would love to hear about yours.

Steve’s report is a look at what is coming next, and is largely, but not only, based on an analysis of what Google is doing, what they are planning on doing, and the emerging ecosystem they are creating. This is fascinating stuff. Steve has recently launched a must-read blog, Beyond Search, where you can get a peek at some of what will be in our report. For example, see his thoughts on enterprise search terminology.

Both reports will be important tools for enterprise IT strategists and executives. We’ll keep you posted on their progress.

Search Adoption is a Tricky Business: Knowledge Needed

January 24, 2008 / Lynda Moulton / 0 Comments

Enterprise search applications abound in the technology marketplace, from embedded search to specialized e-discovery solutions to search engines for crawling and indexing the entire intranet of an organization. So, why is there so much dissatisfaction with results and heaps of stories of buyer’s remorse? Are we on the cusp of a new wave of semantic search options or better ways to federate our universe of content within and outside the enterprise? Who are the experts on enterprise search anyway?

You might read this blog because you know me from the knowledge management (KM) arena, or from my past life as the founder of an integrated enterprise library automation company. In the KM world a recurring theme is the need to leverage expertise, best done in an environment where it is easy to connect with the experts but that seems to be a dim option in many enterprises. In the corporate library world the intent is to aggregate and filter a substantive domain of content, expertise and knowledge assets on behalf of the specialized interests of the enterprise, too often a legacy model of enterprise infrastructure. Librarians have long been innovators at adopting and leveraging advanced technologies but they have also been a concentrating force for facilitating shared expertise. In fact, special librarians excel at providing access to experts.

We are drowning in technological options, not the least of which is enterprise search and its complexity of feature laden choices. However, it is darned hard to find instances of full search tool adoption or users who love the search tools they are delivered on their intranets. So, I am adopting my KM and library science modes to elevate the discussion about search to a decidedly non-technical conversation.

I really want to learn what you know about enterprise search, what you have learned, discovered and experienced over the past two or three years. This blog and the work I do with The Gilbane Group is about getting readers to the best and most appropriate search solutions that can make positive contributions in their enterprises. Knowing who is using what and where it has succeeded or what problems and issues were encountered is information I can use to communicate, in aggregate, those experiences. I am reaching out to you and those you refer to complete a five minute survey to open the door to more discussion. Please use this link to participate right now Click Here to take survey. You will then have the option to get the resulting details in my upcoming research study on enterprise search.

Just to prove that I still follow exciting technologies, as well, I want to relay a couple of new items. First is a recent category in search, “active intelligence,” adopted as Attivio’s tag line. This is a start-up led by Ali Riaz and officially launched this week from Newton, MA. Then, to get a steady feed of all things enterprise search from guru Steve Arnold, check out his new blog, a lead up to the forthcoming Beyond Search: What to Do When Your Search Engine Doesn’t Work to be published by The Gilbane Group. You’ll be transported from the historical, to the here and now, to the newest tools on his radar screen as you page from one blog entry to another.

Nothing Like a Move by Microsoft to Stir up Analysis and Expectations

January 16, 2008 / Lynda Moulton / 0 Comments

Since I weighed in last week on the Microsoft acquisition of FAST Search & Transfer, I have probably read 50+ blog entries and articles on the event. I have also talked to other analysts, received emails from numerous search vendors summarizing their thoughts and expectations about the enterprise search market and had a fair number of phone calls asking questions about what it means. The questions range from “Did Microsoft pay too much?, to “Please define enterprise search,” to “What are the next acquisitions in this market going to be going to be?” My short and flippant answers would be “No,” “Do you have a few hours?” and “Everyone and no one.”

I have seen some excellent analysis contributing relevant commentary to this discussion, some misinterpretation of what the distinction’s are between enterprise search and Web search, and some conclusions that I would seriously debate. You’ll forgive me if I don’t include links to the pieces that influenced the following comments. But one by Curt Monash in his piece on January 14 summarized the state of this industry and its already long history. It is noteworthy that while the popular technology press has only recently begun to write about enterprise search, it has been around for decades in different forms and in a short piece he manages to capture the highlights and current state.

Other commentary seems to imply that Microsoft is not really positioning itself to compete with Google because Google is really about Web (Internet) searching and Microsoft is not. This implies that FAST has no understanding of Web searching. Several points must be made:

FAST Search & Transfer has been involved in many aspects of search technologies for a decade. Soon after landing on our shores it was the search engine of choice for the U.S. government’s unifying search engine to support Internet-based searching of agency Web sites by the public. Since then it has helped countless enterprises (e.g. governments, manufacturers, e-commerce companies) expose their content, products and services via the Web. FAST knows a lot about how to make Web search better for all kinds of applications and they will bring that expertise to Microsoft.
Google is exploiting the Web to deliver free business software tools that directly challenge Microsoft stronghold ( e.g. email, word processing). This will not go unanswered by the largest supplier of office automation software.
Google has several thousand Google Enterprise Search Appliances installed in all types of enterprises around the world, so it is already as widely deployed in enterprises in terms of numbers as FAST, albeit at much lower prices and for simpler application. That doesn’t mean that they are not satisfying a very practical need for a lot of organizations where it is “good enough.”

For more on the competition between the two check this article out.

Enterprise search has been implied to mean only search across all content for an entire enterprise. This raises another fundamental problem of perception. Basically, there are few to no instances of a single enterprise search engine being the only search solution for any major enterprise. Even when an organization “standardizes” on one product for its enterprise search, there will be dozens of other instances of search deployed for groups, divisions, and embedded within applications. Just two examples are the use of Vivisimo now used for USA.gov to give the public access to government agency public content, even as each agency uses different search engines for internal use. Also, there is IBM, which offers the OmniFind suite of enterprise search products, but uses Endeca internally for its Global Services Business enterprise.

Finally, on the issue of expectations, most of the vendors I have heard from are excited that the Microsoft announcement confirms the existence of an enterprise search market. They know that revenues for enterprise search, compared to Web search, have been miniscule. But now that Microsoft is investing heavily in it, they hope that top management across all industries will see it as a software solution to procure. Many analysts are expecting other major acquisitions, perhaps soon. Frequently mentioned buyers are Oracle and IBM but both have already made major acquisitions of search and content products, and both already offer enterprise search solutions. It is going to be quite some time before Microsoft sorts out all the pieces of FAST IP and decides how to package them. Other market acquisitions will surely come. The question is whether the next to be acquired will be large search companies with complex and expensive offerings bought by major software corporations. Or will search products targeting specific enterprise search markets be a better buy to make an immediate impact for companies seeking broader presence in enterprise search as a complementary offering to other tools. There are a lot of enterprise search problems to be solved and a lot of players to divvy up the evolving business for a while to come.

W3C Opens Data on the Web with SPARQL

January 15, 2008 / NewsShark

W3C (The World Wide Web Consortium) announced the publication of SPARQL, the key standard for opening up data on the Semantic Web. With SPARQL query technology, pronounced “sparkle,” people can focus on what they want to know rather than on the database technology or data format used behind the scenes to store the data. Because SPARQL queries express high-level goals, it is easier to extend them to unanticipated data sources, or even to port them to new applications. Many successful query languages exist, including standards such as SQL and XQuery. These were primarily designed for queries limited to a single product, format, type of information, or local data store. Traditionally, it has been necessary to formulate the same high-level query differently depending on application or the specific arrangement chosen for the relational database. And when querying multiple data sources it has been necessary to write logic to merge the results. These limitations have imposed higher developer costs and created barriers to incorporating new data sources. The goal of the Semantic Web is to enable people to share, merge, and reuse data globally. SPARQL is designed for use at the scale of the Web, and thus enables queries over distributed data sources, independent of format. Because SPARQL has no tie to a specific database format, it can be used to take advantage of “Web 2.0” data and mash it up with other Semantic Web resources. Furthermore, because disparate data sources may not have the same ‘shape’ or share the same properties, SPARQL is designed to query non-uniform data. The SPARQL specification defines a query language and a protocol and works with the other core Semantic Web technologies from W3C: Resource Description Framework (RDF) for representing data; RDF Schema; Web Ontology Language (OWL) for building vocabularies; and Gleaning Resource Descriptions from Dialects of Languages (GRDDL), for automatically extracting Semantic Web data from documents. SPARQL also makes use of other W3C standards found in Web services implementations, such as Web Services Description Language (WSDL). http://www.w3.org/

Microsoft and FAST

January 9, 2008 / Frank Gilbane / 0 Comments

Yesterday was obviously a big day in the enterprise search space. “Enterprise search”, as opposed to web search news, doesn’t usually make the New York Times, Wall Street Journal Boston Globe etc. We (especially Lynda!) spent a lot of time yesterday just dealing with all the inquiries. Lynda posted her initial thoughts before the analyst call yesterday, as did Steve Arnold. Both will certainly have more to say. In addition to their blogging keep an eye out for the two reports we’ll be publishing this Spring: Enterprise Search Markets and Applications: Capitalizing on Emerging Demand, by Lynda Moulton, and Beyond Search: What to do When You’re Enterprise Search System Doesn’t Work, by Steve Arnold.

Category: Semantic technologies (Page 34 of 72)

Taxonomy and Enterprise Search

Search Engines Under the Hood

Search Behind the Firewall aka Enterprise Search

Beyond Search and Search

Search Adoption is a Tricky Business: Knowledge Needed

Nothing Like a Move by Microsoft to Stir up Analysis and Expectations

W3C Opens Data on the Web with SPARQL

Microsoft and FAST

Subscribe to the Gilbane Advisor

Choose Language

Topics we cover

Policies

Contact