Search (51 results, page 1 of 3)

Wiegmann, S.: Hättest du die Titanic überlebt? : Eine kurze Einführung in das Data Mining mit freier Software (2023) 0.04

0.041977502 = product of:
  0.139925 = sum of:
    0.034665532 = weight(_text_:software in 876) [ClassicSimilarity], result of:
      0.034665532 = score(doc=876,freq=4.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.43390724 = fieldWeight in 876, product of:
          2.0 = tf(freq=4.0), with freq of:
            4.0 = termFreq=4.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.0546875 = fieldNorm(doc=876)
    0.015301661 = weight(_text_:und in 876) [ClassicSimilarity], result of:
      0.015301661 = score(doc=876,freq=8.0), product of:
        0.044633795 = queryWeight, product of:
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.02013827 = queryNorm
        0.34282678 = fieldWeight in 876, product of:
          2.828427 = tf(freq=8.0), with freq of:
            8.0 = termFreq=8.0
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.0546875 = fieldNorm(doc=876)
    0.034665532 = weight(_text_:software in 876) [ClassicSimilarity], result of:
      0.034665532 = score(doc=876,freq=4.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.43390724 = fieldWeight in 876, product of:
          2.0 = tf(freq=4.0), with freq of:
            4.0 = termFreq=4.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.0546875 = fieldNorm(doc=876)
    0.0109904595 = weight(_text_:der in 876) [ClassicSimilarity], result of:
      0.0109904595 = score(doc=876,freq=4.0), product of:
        0.044984195 = queryWeight, product of:
          2.2337668 = idf(docFreq=12875, maxDocs=44218)
          0.02013827 = queryNorm
        0.24431825 = fieldWeight in 876, product of:
          2.0 = tf(freq=4.0), with freq of:
            4.0 = termFreq=4.0
          2.2337668 = idf(docFreq=12875, maxDocs=44218)
          0.0546875 = fieldNorm(doc=876)
    0.009636286 = product of:
      0.019272571 = sum of:
        0.019272571 = weight(_text_:29 in 876) [ClassicSimilarity], result of:
          0.019272571 = score(doc=876,freq=2.0), product of:
            0.070840135 = queryWeight, product of:
              3.5176873 = idf(docFreq=3565, maxDocs=44218)
              0.02013827 = queryNorm
            0.27205724 = fieldWeight in 876, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              3.5176873 = idf(docFreq=3565, maxDocs=44218)
              0.0546875 = fieldNorm(doc=876)
      0.5 = coord(1/2)
    0.034665532 = weight(_text_:software in 876) [ClassicSimilarity], result of:
      0.034665532 = score(doc=876,freq=4.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.43390724 = fieldWeight in 876, product of:
          2.0 = tf(freq=4.0), with freq of:
            4.0 = termFreq=4.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.0546875 = fieldNorm(doc=876)
  0.3 = coord(6/20)

Abstract: Am 10. April 1912 ging Elisabeth Walton Allen an Bord der "Titanic", um ihr Hab und Gut nach England zu holen. Eines Nachts wurde sie von ihrer aufgelösten Tante geweckt, deren Kajüte unter Wasser stand. Wie steht es um Elisabeths Chancen und hätte man selbst das Unglück damals überlebt? Das Titanic-Orakel ist eine algorithmusbasierte App, die entsprechende Prognosen aufstellt und im Rahmen des Kurses "Data Science" am Department Information der HAW Hamburg entstanden ist. Dieser Beitrag zeigt Schritt für Schritt, wie die App unter Verwendung freier Software entwickelt wurde. Code und Daten werden zur Nachnutzung bereitgestellt.
Date: 28. 1.2022 11:05:29

Leydesdorff, L.; Persson, O.: Mapping the geography of science : distribution patterns and networks of relations among cities and institutes (2010) 0.03

0.034343183 = product of:
  0.11447728 = sum of:
    0.017148608 = weight(_text_:23 in 3704) [ClassicSimilarity], result of:
      0.017148608 = score(doc=3704,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.23759183 = fieldWeight in 3704, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.046875 = fieldNorm(doc=3704)
    0.017148608 = weight(_text_:23 in 3704) [ClassicSimilarity], result of:
      0.017148608 = score(doc=3704,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.23759183 = fieldWeight in 3704, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.046875 = fieldNorm(doc=3704)
    0.021010485 = weight(_text_:software in 3704) [ClassicSimilarity], result of:
      0.021010485 = score(doc=3704,freq=2.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.2629875 = fieldWeight in 3704, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.046875 = fieldNorm(doc=3704)
    0.017148608 = weight(_text_:23 in 3704) [ClassicSimilarity], result of:
      0.017148608 = score(doc=3704,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.23759183 = fieldWeight in 3704, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.046875 = fieldNorm(doc=3704)
    0.021010485 = weight(_text_:software in 3704) [ClassicSimilarity], result of:
      0.021010485 = score(doc=3704,freq=2.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.2629875 = fieldWeight in 3704, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.046875 = fieldNorm(doc=3704)
    0.021010485 = weight(_text_:software in 3704) [ClassicSimilarity], result of:
      0.021010485 = score(doc=3704,freq=2.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.2629875 = fieldWeight in 3704, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.046875 = fieldNorm(doc=3704)
  0.3 = coord(6/20)

Abstract: Using Google Earth, Google Maps, and/or network visualization programs such as Pajek, one can overlay the network of relations among addresses in scientific publications onto the geographic map. The authors discuss the pros and cons of various options, and provide software (freeware) for bridging existing gaps between the Science Citation Indices (Thomson Reuters) and Scopus (Elsevier), on the one hand, and these various visualization tools on the other. At the level of city names, the global map can be drawn reliably on the basis of the available address information. At the level of the names of organizations and institutes, there are problems of unification both in the ISI databases and with Scopus. Pajek enables a combination of visualization and statistical analysis, whereas the Google Maps and its derivatives provide superior tools on the Internet.
Date: 23. 7.2010 13:10:08

Klein, H.: Web Content Mining (2004) 0.03
```
0.029254396 = product of:
  0.11701758 = sum of:
    0.028013978 = weight(_text_:software in 3154) [ClassicSimilarity], result of:
      0.028013978 = score(doc=3154,freq=8.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.35064998 = fieldWeight in 3154, product of:
          2.828427 = tf(freq=8.0), with freq of:
            8.0 = termFreq=8.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.03125 = fieldNorm(doc=3154)
    0.013115709 = weight(_text_:und in 3154) [ClassicSimilarity], result of:
      0.013115709 = score(doc=3154,freq=18.0), product of:
        0.044633795 = queryWeight, product of:
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.02013827 = queryNorm
        0.29385152 = fieldWeight in 3154, product of:
          4.2426405 = tf(freq=18.0), with freq of:
            18.0 = termFreq=18.0
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.03125 = fieldNorm(doc=3154)
    0.028013978 = weight(_text_:software in 3154) [ClassicSimilarity], result of:
      0.028013978 = score(doc=3154,freq=8.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.35064998 = fieldWeight in 3154, product of:
          2.828427 = tf(freq=8.0), with freq of:
            8.0 = termFreq=8.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.03125 = fieldNorm(doc=3154)
    0.019859934 = weight(_text_:der in 3154) [ClassicSimilarity], result of:
      0.019859934 = score(doc=3154,freq=40.0), product of:
        0.044984195 = queryWeight, product of:
          2.2337668 = idf(docFreq=12875, maxDocs=44218)
          0.02013827 = queryNorm
        0.44148692 = fieldWeight in 3154, product of:
          6.3245554 = tf(freq=40.0), with freq of:
            40.0 = termFreq=40.0
          2.2337668 = idf(docFreq=12875, maxDocs=44218)
          0.03125 = fieldNorm(doc=3154)
    0.028013978 = weight(_text_:software in 3154) [ClassicSimilarity], result of:
      0.028013978 = score(doc=3154,freq=8.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.35064998 = fieldWeight in 3154, product of:
          2.828427 = tf(freq=8.0), with freq of:
            8.0 = termFreq=8.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.03125 = fieldNorm(doc=3154)
  0.25 = coord(5/20)
```
Abstract

Web Mining - ein Schlagwort, das mit der Verbreitung des Internets immer öfter zu lesen und zu hören ist. Die gegenwärtige Forschung beschäftigt sich aber eher mit dem Nutzungsverhalten der Internetnutzer, und ein Blick in Tagungsprogramme einschlägiger Konferenzen (z.B. GOR - German Online Research) zeigt, dass die Analyse der Inhalte kaum Thema ist. Auf der GOR wurden 1999 zwei Vorträge zu diesem Thema gehalten, auf der Folgekonferenz 2001 kein einziger. Web Mining ist der Oberbegriff für zwei Typen von Web Mining: Web Usage Mining und Web Content Mining. Unter Web Usage Mining versteht man das Analysieren von Daten, wie sie bei der Nutzung des WWW anfallen und von den Servern protokolliert wenden. Man kann ermitteln, welche Seiten wie oft aufgerufen wurden, wie lange auf den Seiten verweilt wurde und vieles andere mehr. Beim Web Content Mining wird der Inhalt der Webseiten untersucht, der nicht nur Text, sondern auf Bilder, Video- und Audioinhalte enthalten kann. Die Software für die Analyse von Webseiten ist in den Grundzügen vorhanden, doch müssen die meisten Webseiten für die entsprechende Analysesoftware erst aufbereitet werden. Zuerst müssen die relevanten Websites ermittelt werden, die die gesuchten Inhalte enthalten. Das geschieht meist mit Suchmaschinen, von denen es mittlerweile Hunderte gibt. Allerdings kann man nicht davon ausgehen, dass die Suchmaschinen alle existierende Webseiten erfassen. Das ist unmöglich, denn durch das schnelle Wachstum des Internets kommen täglich Tausende von Webseiten hinzu, und bereits bestehende ändern sich der werden gelöscht. Oft weiß man auch nicht, wie die Suchmaschinen arbeiten, denn das gehört zu den Geschäftsgeheimnissen der Betreiber. Man muss also davon ausgehen, dass die Suchmaschinen nicht alle relevanten Websites finden (können). Der nächste Schritt ist das Herunterladen der Websites, dafür gibt es Software, die unter den Bezeichnungen OfflineReader oder Webspider zu finden ist. Das Ziel dieser Programme ist, die Website in einer Form herunterzuladen, die es erlaubt, die Website offline zu betrachten. Die Struktur der Website wird in der Regel beibehalten. Wer die Inhalte einer Website analysieren will, muss also alle Dateien mit seiner Analysesoftware verarbeiten können. Software für Inhaltsanalyse geht davon aus, dass nur Textinformationen in einer einzigen Datei verarbeitet werden. QDA Software (qualitative data analysis) verarbeitet dagegen auch Audiound Videoinhalte sowie internetspezifische Kommunikation wie z.B. Chats.

Series

Fortschritte in der Wissensorganisation; Bd.7

Source

Wissensorganisation und Edutainment: Wissen im Spannungsfeld von Gesellschaft, Gestaltung und Industrie. Proceedings der 7. Tagung der Deutschen Sektion der Internationalen Gesellschaft für Wissensorganisation, Berlin, 21.-23.3.2001. Hrsg.: C. Lehner, H.P. Ohly u. G. Rahmstorf

Brückner, T.; Dambeck, H.: Sortierautomaten : Grundlagen der Textklassifizierung (2003) 0.03

0.027017068 = product of:
  0.10806827 = sum of:
    0.028013978 = weight(_text_:software in 2398) [ClassicSimilarity], result of:
      0.028013978 = score(doc=2398,freq=2.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.35064998 = fieldWeight in 2398, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.0625 = fieldNorm(doc=2398)
    0.015144716 = weight(_text_:und in 2398) [ClassicSimilarity], result of:
      0.015144716 = score(doc=2398,freq=6.0), product of:
        0.044633795 = queryWeight, product of:
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.02013827 = queryNorm
        0.33931053 = fieldWeight in 2398, product of:
          2.4494898 = tf(freq=6.0), with freq of:
            6.0 = termFreq=6.0
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.0625 = fieldNorm(doc=2398)
    0.028013978 = weight(_text_:software in 2398) [ClassicSimilarity], result of:
      0.028013978 = score(doc=2398,freq=2.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.35064998 = fieldWeight in 2398, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.0625 = fieldNorm(doc=2398)
    0.008881632 = weight(_text_:der in 2398) [ClassicSimilarity], result of:
      0.008881632 = score(doc=2398,freq=2.0), product of:
        0.044984195 = queryWeight, product of:
          2.2337668 = idf(docFreq=12875, maxDocs=44218)
          0.02013827 = queryNorm
        0.19743896 = fieldWeight in 2398, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          2.2337668 = idf(docFreq=12875, maxDocs=44218)
          0.0625 = fieldNorm(doc=2398)
    0.028013978 = weight(_text_:software in 2398) [ClassicSimilarity], result of:
      0.028013978 = score(doc=2398,freq=2.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.35064998 = fieldWeight in 2398, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.0625 = fieldNorm(doc=2398)
  0.25 = coord(5/20)

Abstract: Rechnung, Kündigung oder Adressänderung? Eingehende Briefe und E-Mails werden immer häufiger von Software statt aufwändig von Menschenhand sortiert. Die Textklassifizierer arbeiten erstaunlich genau. Sie fahnden auch nach ähnlichen Texten und sorgen so für einen schnellen Überblick. Ihre Werkzeuge sind Linguistik, Statistik und Logik

Loonus, Y.: Einsatzbereiche der KI und ihre Relevanz für Information Professionals (2017) 0.02

0.023104325 = product of:
  0.0924173 = sum of:
    0.021010485 = weight(_text_:software in 5668) [ClassicSimilarity], result of:
      0.021010485 = score(doc=5668,freq=2.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.2629875 = fieldWeight in 5668, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.046875 = fieldNorm(doc=5668)
    0.016063396 = weight(_text_:und in 5668) [ClassicSimilarity], result of:
      0.016063396 = score(doc=5668,freq=12.0), product of:
        0.044633795 = queryWeight, product of:
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.02013827 = queryNorm
        0.35989314 = fieldWeight in 5668, product of:
          3.4641016 = tf(freq=12.0), with freq of:
            12.0 = termFreq=12.0
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.046875 = fieldNorm(doc=5668)
    0.021010485 = weight(_text_:software in 5668) [ClassicSimilarity], result of:
      0.021010485 = score(doc=5668,freq=2.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.2629875 = fieldWeight in 5668, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.046875 = fieldNorm(doc=5668)
    0.013322448 = weight(_text_:der in 5668) [ClassicSimilarity], result of:
      0.013322448 = score(doc=5668,freq=8.0), product of:
        0.044984195 = queryWeight, product of:
          2.2337668 = idf(docFreq=12875, maxDocs=44218)
          0.02013827 = queryNorm
        0.29615843 = fieldWeight in 5668, product of:
          2.828427 = tf(freq=8.0), with freq of:
            8.0 = termFreq=8.0
          2.2337668 = idf(docFreq=12875, maxDocs=44218)
          0.046875 = fieldNorm(doc=5668)
    0.021010485 = weight(_text_:software in 5668) [ClassicSimilarity], result of:
      0.021010485 = score(doc=5668,freq=2.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.2629875 = fieldWeight in 5668, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.046875 = fieldNorm(doc=5668)
  0.25 = coord(5/20)

Abstract: Es liegt in der Natur des Menschen, Erfahrungen und Ideen in Wort und Schrift mit anderen teilen zu wollen. So produzieren wir jeden Tag gigantische Mengen an Texten, die in digitaler Form geteilt und abgelegt werden. The Radicati Group schätzt, dass 2017 täglich 269 Milliarden E-Mails versendet und empfangen werden. Hinzu kommen größtenteils unstrukturierte Daten wie Soziale Medien, Presse, Websites und firmeninterne Systeme, beispielsweise in Form von CRM-Software oder PDF-Dokumenten. Der weltweite Bestand an unstrukturierten Daten wächst so rasant, dass es kaum möglich ist, seinen Umfang zu quantifizieren. Der Versuch, eine belastbare Zahl zu recherchieren, führt unweigerlich zu diversen Artikeln, die den Anteil unstrukturierter Texte am gesamten Datenbestand auf 80% schätzen. Auch wenn nicht mehr einwandfrei nachvollziehbar ist, woher diese Zahl stammt, kann bei kritischer Reflexion unseres Tagesablaufs kaum bezweifelt werden, dass diese Daten von großer wirtschaftlicher Relevanz sind.

Howlett, D.: Digging deep for treasure (1998) 0.02

0.020578332 = product of:
  0.13718888 = sum of:
    0.045729626 = weight(_text_:23 in 4544) [ClassicSimilarity], result of:
      0.045729626 = score(doc=4544,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.63357824 = fieldWeight in 4544, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.125 = fieldNorm(doc=4544)
    0.045729626 = weight(_text_:23 in 4544) [ClassicSimilarity], result of:
      0.045729626 = score(doc=4544,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.63357824 = fieldWeight in 4544, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.125 = fieldNorm(doc=4544)
    0.045729626 = weight(_text_:23 in 4544) [ClassicSimilarity], result of:
      0.045729626 = score(doc=4544,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.63357824 = fieldWeight in 4544, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.125 = fieldNorm(doc=4544)
  0.15 = coord(3/20)

Date: 26. 3.2000 16:35:23

Tunbridge, N.: Semiology put to data mining (1999) 0.02

0.020578332 = product of:
  0.13718888 = sum of:
    0.045729626 = weight(_text_:23 in 6782) [ClassicSimilarity], result of:
      0.045729626 = score(doc=6782,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.63357824 = fieldWeight in 6782, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.125 = fieldNorm(doc=6782)
    0.045729626 = weight(_text_:23 in 6782) [ClassicSimilarity], result of:
      0.045729626 = score(doc=6782,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.63357824 = fieldWeight in 6782, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.125 = fieldNorm(doc=6782)
    0.045729626 = weight(_text_:23 in 6782) [ClassicSimilarity], result of:
      0.045729626 = score(doc=6782,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.63357824 = fieldWeight in 6782, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.125 = fieldNorm(doc=6782)
  0.15 = coord(3/20)

Source: Online and CD-ROM review. 23(1999) no.5, S.303-305

Heyer, G.; Läuter, M.; Quasthoff, U.; Wolff, C.: Texttechnologische Anwendungen am Beispiel Text Mining (2000) 0.02

0.019024778 = product of:
  0.07609911 = sum of:
    0.017148608 = weight(_text_:23 in 5565) [ClassicSimilarity], result of:
      0.017148608 = score(doc=5565,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.23759183 = fieldWeight in 5565, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.046875 = fieldNorm(doc=5565)
    0.017148608 = weight(_text_:23 in 5565) [ClassicSimilarity], result of:
      0.017148608 = score(doc=5565,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.23759183 = fieldWeight in 5565, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.046875 = fieldNorm(doc=5565)
    0.013115709 = weight(_text_:und in 5565) [ClassicSimilarity], result of:
      0.013115709 = score(doc=5565,freq=8.0), product of:
        0.044633795 = queryWeight, product of:
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.02013827 = queryNorm
        0.29385152 = fieldWeight in 5565, product of:
          2.828427 = tf(freq=8.0), with freq of:
            8.0 = termFreq=8.0
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.046875 = fieldNorm(doc=5565)
    0.017148608 = weight(_text_:23 in 5565) [ClassicSimilarity], result of:
      0.017148608 = score(doc=5565,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.23759183 = fieldWeight in 5565, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.046875 = fieldNorm(doc=5565)
    0.011537581 = weight(_text_:der in 5565) [ClassicSimilarity], result of:
      0.011537581 = score(doc=5565,freq=6.0), product of:
        0.044984195 = queryWeight, product of:
          2.2337668 = idf(docFreq=12875, maxDocs=44218)
          0.02013827 = queryNorm
        0.25648075 = fieldWeight in 5565, product of:
          2.4494898 = tf(freq=6.0), with freq of:
            6.0 = termFreq=6.0
          2.2337668 = idf(docFreq=12875, maxDocs=44218)
          0.046875 = fieldNorm(doc=5565)
  0.25 = coord(5/20)

Abstract: Die zunehmende Menge von Informationen und deren weltweite Verfügbarkeit auf der Basis moderner Internet Technologie machen es erforderlich, Informationen nach inhaltlichen Kriterien zu strukturieren und zu bewerten sowie nach inhaltlichen Kriterien weiter zu verarbeiten. Vom Standpunkt des Benutzers aus sind dabei folgende Fälle zu unterscheiden: Handelt es sich bei den gesuchten Informationen um strukturierle Daten (z.B. in einer SQL-Datenbank) oder unstrukturierte Daten (z.B. grosse Texte)? Ist bekannt, welche Daten benötigt werden und wie sie zu finden sind? Oder ist vor dein Zugriff auf die Daten noch nicht bekannt welche Ergebnisse erwartet werden?
Source: Sprachtechnologie für eine dynamische Wirtschaft im Medienzeitalter - Language technologies for dynamic business in the age of the media - L'ingénierie linguistique au service de la dynamisation économique à l'ère du multimédia: Tagungsakten der XXVI. Jahrestagung der Internationalen Vereinigung Sprache und Wirtschaft e.V., 23.-25.11.2000, Fachhochschule Köln. Hrsg.: K.-D. Schmitz

Kraker, P.; Kittel, C,; Enkhbayar, A.: Open Knowledge Maps : creating a visual interface to the world's scientific knowledge based on natural language processing (2016) 0.01

0.013917861 = product of:
  0.0695893 = sum of:
    0.021010485 = weight(_text_:software in 3205) [ClassicSimilarity], result of:
      0.021010485 = score(doc=3205,freq=2.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.2629875 = fieldWeight in 3205, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.046875 = fieldNorm(doc=3205)
    0.0065578544 = weight(_text_:und in 3205) [ClassicSimilarity], result of:
      0.0065578544 = score(doc=3205,freq=2.0), product of:
        0.044633795 = queryWeight, product of:
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.02013827 = queryNorm
        0.14692576 = fieldWeight in 3205, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.046875 = fieldNorm(doc=3205)
    0.021010485 = weight(_text_:software in 3205) [ClassicSimilarity], result of:
      0.021010485 = score(doc=3205,freq=2.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.2629875 = fieldWeight in 3205, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.046875 = fieldNorm(doc=3205)
    0.021010485 = weight(_text_:software in 3205) [ClassicSimilarity], result of:
      0.021010485 = score(doc=3205,freq=2.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.2629875 = fieldWeight in 3205, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.046875 = fieldNorm(doc=3205)
  0.2 = coord(4/20)

Abstract: The goal of Open Knowledge Maps is to create a visual interface to the world's scientific knowledge. The base for this visual interface consists of so-called knowledge maps, which enable the exploration of existing knowledge and the discovery of new knowledge. Our open source knowledge mapping software applies a mixture of summarization techniques and similarity measures on article metadata, which are iteratively chained together. After processing, the representation is saved in a database for use in a web visualization. In the future, we want to create a space for collective knowledge mapping that brings together individuals and communities involved in exploration and discovery. We want to enable people to guide each other in their discovery by collaboratively annotating and modifying the automatically created maps.
Content: Beitrag in einem Themenschwerpunkt 'Computerlinguistik und Bibliotheken'. Vgl.: http://0277.ch/ojs/index.php/cdrs_0277/article/view/157/355.

Drees, B.: Text und data mining : Herausforderungen und Möglichkeiten für Bibliotheken (2016) 0.01

0.012446469 = product of:
  0.08297645 = sum of:
    0.020737756 = weight(_text_:und in 3952) [ClassicSimilarity], result of:
      0.020737756 = score(doc=3952,freq=20.0), product of:
        0.044633795 = queryWeight, product of:
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.02013827 = queryNorm
        0.46462005 = fieldWeight in 3952, product of:
          4.472136 = tf(freq=20.0), with freq of:
            20.0 = termFreq=20.0
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.046875 = fieldNorm(doc=3952)
    0.050701115 = weight(_text_:methoden in 3952) [ClassicSimilarity], result of:
      0.050701115 = score(doc=3952,freq=4.0), product of:
        0.10436003 = queryWeight, product of:
          5.1821747 = idf(docFreq=674, maxDocs=44218)
          0.02013827 = queryNorm
        0.48582888 = fieldWeight in 3952, product of:
          2.0 = tf(freq=4.0), with freq of:
            4.0 = termFreq=4.0
          5.1821747 = idf(docFreq=674, maxDocs=44218)
          0.046875 = fieldNorm(doc=3952)
    0.011537581 = weight(_text_:der in 3952) [ClassicSimilarity], result of:
      0.011537581 = score(doc=3952,freq=6.0), product of:
        0.044984195 = queryWeight, product of:
          2.2337668 = idf(docFreq=12875, maxDocs=44218)
          0.02013827 = queryNorm
        0.25648075 = fieldWeight in 3952, product of:
          2.4494898 = tf(freq=6.0), with freq of:
            6.0 = termFreq=6.0
          2.2337668 = idf(docFreq=12875, maxDocs=44218)
          0.046875 = fieldNorm(doc=3952)
  0.15 = coord(3/20)

Abstract: Text und Data Mining (TDM) gewinnt als wissenschaftliche Methode zunehmend an Bedeutung und stellt wissenschaftliche Bibliotheken damit vor neue Herausforderungen, bietet gleichzeitig aber auch neue Möglichkeiten. Der vorliegende Beitrag gibt einen Überblick über das Thema TDM aus bibliothekarischer Sicht. Hierzu wird der Begriff Text und Data Mining im Kontext verwandter Begriffe diskutiert sowie Ziele, Aufgaben und Methoden von TDM erläutert. Diese werden anhand beispielhafter TDM-Anwendungen in Wissenschaft und Forschung illustriert. Ferner werden technische und rechtliche Probleme und Hindernisse im TDM-Kontext dargelegt. Abschließend wird die Relevanz von TDM für Bibliotheken, sowohl in ihrer Rolle als Informationsvermittler und -anbieter als auch als Anwender von TDM-Methoden, aufgezeigt. Zudem wurde im Rahmen dieser Arbeit eine Befragung der Betreiber von Dokumentenservern an Bibliotheken in Deutschland zum aktuellen Umgang mit TDM durchgeführt, die zeigt, dass hier noch viel Ausbaupotential besteht. Die dem Artikel zugrunde liegenden Forschungsdaten sind unter dem DOI 10.11588/data/10090 publiziert.

Peters, G.; Gaese, V.: ¬Das DocCat-System in der Textdokumentation von G+J (2003) 0.01
```
0.011678714 = product of:
  0.05839357 = sum of:
    0.015144716 = weight(_text_:und in 1507) [ClassicSimilarity], result of:
      0.015144716 = score(doc=1507,freq=24.0), product of:
        0.044633795 = queryWeight, product of:
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.02013827 = queryNorm
        0.33931053 = fieldWeight in 1507, product of:
          4.8989797 = tf(freq=24.0), with freq of:
            24.0 = termFreq=24.0
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.03125 = fieldNorm(doc=1507)
    0.0117492955 = weight(_text_:der in 1507) [ClassicSimilarity], result of:
      0.0117492955 = score(doc=1507,freq=14.0), product of:
        0.044984195 = queryWeight, product of:
          2.2337668 = idf(docFreq=12875, maxDocs=44218)
          0.02013827 = queryNorm
        0.2611872 = fieldWeight in 1507, product of:
          3.7416575 = tf(freq=14.0), with freq of:
            14.0 = termFreq=14.0
          2.2337668 = idf(docFreq=12875, maxDocs=44218)
          0.03125 = fieldNorm(doc=1507)
    0.027861618 = product of:
      0.055723235 = sum of:
        0.055723235 = weight(_text_:programmierung in 1507) [ClassicSimilarity], result of:
          0.055723235 = score(doc=1507,freq=2.0), product of:
            0.15934804 = queryWeight, product of:
              7.912698 = idf(docFreq=43, maxDocs=44218)
              0.02013827 = queryNorm
            0.34969515 = fieldWeight in 1507, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              7.912698 = idf(docFreq=43, maxDocs=44218)
              0.03125 = fieldNorm(doc=1507)
      0.5 = coord(1/2)
    0.0036379434 = product of:
      0.01091383 = sum of:
        0.01091383 = weight(_text_:22 in 1507) [ClassicSimilarity], result of:
          0.01091383 = score(doc=1507,freq=2.0), product of:
            0.07052079 = queryWeight, product of:
              3.5018296 = idf(docFreq=3622, maxDocs=44218)
              0.02013827 = queryNorm
            0.15476047 = fieldWeight in 1507, product of:
              1.4142135 = tf(freq=2.0), with freq of:
                2.0 = termFreq=2.0
              3.5018296 = idf(docFreq=3622, maxDocs=44218)
              0.03125 = fieldNorm(doc=1507)
      0.33333334 = coord(1/3)
  0.2 = coord(4/20)
```
Abstract

Wir werden einmal die Grundlagen des Text-Mining-Systems bei IBM darstellen, dann werden wir das Projekt etwas umfangreicher und deutlicher darstellen, da kennen wir uns aus. Von daher haben wir zwei Teile, einmal Heidelberg, einmal Hamburg. Noch einmal zur Technologie. Text-Mining ist eine von IBM entwickelte Technologie, die in einer besonderen Ausformung und Programmierung für uns zusammengestellt wurde. Das Projekt hieß bei uns lange Zeit DocText Miner und heißt seit einiger Zeit auf Vorschlag von IBM DocCat, das soll eine Abkürzung für Document-Categoriser sein, sie ist ja auch nett und anschaulich. Wir fangen an mit Text-Mining, das bei IBM in Heidelberg entwickelt wurde. Die verstehen darunter das automatische Indexieren als eine Instanz, also einen Teil von Text-Mining. Probleme werden dabei gezeigt, und das Text-Mining ist eben eine Methode zur Strukturierung von und der Suche in großen Dokumentenmengen, die Extraktion von Informationen und, das ist der hohe Anspruch, von impliziten Zusammenhängen. Das letztere sei dahingestellt. IBM macht das quantitativ, empirisch, approximativ und schnell. das muss man wirklich sagen. Das Ziel, und das ist ganz wichtig für unser Projekt gewesen, ist nicht, den Text zu verstehen, sondern das Ergebnis dieser Verfahren ist, was sie auf Neudeutsch a bundle of words, a bag of words nennen, also eine Menge von bedeutungstragenden Begriffen aus einem Text zu extrahieren, aufgrund von Algorithmen, also im Wesentlichen aufgrund von Rechenoperationen. Es gibt eine ganze Menge von linguistischen Vorstudien, ein wenig Linguistik ist auch dabei, aber nicht die Grundlage der ganzen Geschichte. Was sie für uns gemacht haben, ist also die Annotierung von Pressetexten für unsere Pressedatenbank. Für diejenigen, die es noch nicht kennen: Gruner + Jahr führt eine Textdokumentation, die eine Datenbank führt, seit Anfang der 70er Jahre, da sind z.Z. etwa 6,5 Millionen Dokumente darin, davon etwas über 1 Million Volltexte ab 1993. Das Prinzip war lange Zeit, dass wir die Dokumente, die in der Datenbank gespeichert waren und sind, verschlagworten und dieses Prinzip haben wir auch dann, als der Volltext eingeführt wurde, in abgespeckter Form weitergeführt. Zu diesen 6,5 Millionen Dokumenten gehören dann eben auch ungefähr 10 Millionen Faksimileseiten, weil wir die Faksimiles auch noch standardmäßig aufheben.

Date

22. 4.2003 11:45:36

Source

Medien-Informationsmanagement: Archivarische, dokumentarische, betriebswirtschaftliche, rechtliche und Berufsbild-Aspekte. Hrsg.: Marianne Englert u.a

Haravu, L.J.; Neelameghan, A.: Text mining and data mining in knowledge organization and discovery : the making of knowledge-based products (2003) 0.01

0.011142492 = product of:
  0.07428328 = sum of:
    0.024761094 = weight(_text_:software in 5653) [ClassicSimilarity], result of:
      0.024761094 = score(doc=5653,freq=4.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.30993375 = fieldWeight in 5653, product of:
          2.0 = tf(freq=4.0), with freq of:
            4.0 = termFreq=4.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.0390625 = fieldNorm(doc=5653)
    0.024761094 = weight(_text_:software in 5653) [ClassicSimilarity], result of:
      0.024761094 = score(doc=5653,freq=4.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.30993375 = fieldWeight in 5653, product of:
          2.0 = tf(freq=4.0), with freq of:
            4.0 = termFreq=4.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.0390625 = fieldNorm(doc=5653)
    0.024761094 = weight(_text_:software in 5653) [ClassicSimilarity], result of:
      0.024761094 = score(doc=5653,freq=4.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.30993375 = fieldWeight in 5653, product of:
          2.0 = tf(freq=4.0), with freq of:
            4.0 = termFreq=4.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.0390625 = fieldNorm(doc=5653)
  0.15 = coord(3/20)

Abstract: Discusses the importance of knowledge organization in the context of the information overload caused by the vast quantities of data and information accessible on internal and external networks of an organization. Defines the characteristics of a knowledge-based product. Elaborates on the techniques and applications of text mining in developing knowledge products. Presents two approaches, as case studies, to the making of knowledge products: (1) steps and processes in the planning, designing and development of a composite multilingual multimedia CD product, with the potential international, inter-cultural end users in view, and (2) application of natural language processing software in text mining. Using a text mining software, it is possible to link concept terms from a processed text to a related thesaurus, glossary, schedules of a classification scheme, and facet structured subject representations. Concludes that the products of text mining and data mining could be made more useful if the features of a faceted scheme for subject classification are incorporated into text mining techniques and products.

Gluck , M.: Multimedia exploratory data analysis for geospatial data mining : the case for augmented seriation (2001) 0.01

0.009454718 = product of:
  0.06303145 = sum of:
    0.021010485 = weight(_text_:software in 5214) [ClassicSimilarity], result of:
      0.021010485 = score(doc=5214,freq=2.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.2629875 = fieldWeight in 5214, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.046875 = fieldNorm(doc=5214)
    0.021010485 = weight(_text_:software in 5214) [ClassicSimilarity], result of:
      0.021010485 = score(doc=5214,freq=2.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.2629875 = fieldWeight in 5214, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.046875 = fieldNorm(doc=5214)
    0.021010485 = weight(_text_:software in 5214) [ClassicSimilarity], result of:
      0.021010485 = score(doc=5214,freq=2.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.2629875 = fieldWeight in 5214, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.046875 = fieldNorm(doc=5214)
  0.15 = coord(3/20)

Abstract: To prevent type-one error, statisticians tend to accept the possibility of type-two error, which leads to the rejection of hypotheses later shown to be true. In both Exploratory Data Analysis and data mining the emphasis is more appropriately on the elimination of type-two error. Thus EDA methods, including its visualization tools may be appropriate for Data Mining. Seriation, creates a matrix of observations and variables, where the cells contain an icon whose size represents its value, and permits the movement of rows and columns in order to visually discern patterns. Augmented Seriation, a method of data mining, adds computer graphics, sound, color, and extra dimensions to the matrix so that the analyst has different modalities for pattern observation. Gluck has developed software for such analysis.

Raghavan, V.V.; Deogun, J.S.; Sever, H.: Knowledge discovery and data mining : introduction (1998) 0.01

0.009003021 = product of:
  0.060020134 = sum of:
    0.02000671 = weight(_text_:23 in 2899) [ClassicSimilarity], result of:
      0.02000671 = score(doc=2899,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.27719048 = fieldWeight in 2899, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.0546875 = fieldNorm(doc=2899)
    0.02000671 = weight(_text_:23 in 2899) [ClassicSimilarity], result of:
      0.02000671 = score(doc=2899,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.27719048 = fieldWeight in 2899, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.0546875 = fieldNorm(doc=2899)
    0.02000671 = weight(_text_:23 in 2899) [ClassicSimilarity], result of:
      0.02000671 = score(doc=2899,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.27719048 = fieldWeight in 2899, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.0546875 = fieldNorm(doc=2899)
  0.15 = coord(3/20)

Date: 7. 2.1999 11:23:06

Seidenfaden, U.: Schürfen in Datenbergen : Data-Mining soll möglichst viel Information zu Tage fördern (2001) 0.01
```
0.00807841 = product of:
  0.053856067 = sum of:
    0.010819908 = weight(_text_:und in 6923) [ClassicSimilarity], result of:
      0.010819908 = score(doc=6923,freq=16.0), product of:
        0.044633795 = queryWeight, product of:
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.02013827 = queryNorm
        0.24241515 = fieldWeight in 6923, product of:
          4.0 = tf(freq=16.0), with freq of:
            16.0 = termFreq=16.0
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.02734375 = fieldNorm(doc=6923)
    0.029575652 = weight(_text_:methoden in 6923) [ClassicSimilarity], result of:
      0.029575652 = score(doc=6923,freq=4.0), product of:
        0.10436003 = queryWeight, product of:
          5.1821747 = idf(docFreq=674, maxDocs=44218)
          0.02013827 = queryNorm
        0.28340018 = fieldWeight in 6923, product of:
          2.0 = tf(freq=4.0), with freq of:
            4.0 = termFreq=4.0
          5.1821747 = idf(docFreq=674, maxDocs=44218)
          0.02734375 = fieldNorm(doc=6923)
    0.0134605095 = weight(_text_:der in 6923) [ClassicSimilarity], result of:
      0.0134605095 = score(doc=6923,freq=24.0), product of:
        0.044984195 = queryWeight, product of:
          2.2337668 = idf(docFreq=12875, maxDocs=44218)
          0.02013827 = queryNorm
        0.29922754 = fieldWeight in 6923, product of:
          4.8989797 = tf(freq=24.0), with freq of:
            24.0 = termFreq=24.0
          2.2337668 = idf(docFreq=12875, maxDocs=44218)
          0.02734375 = fieldNorm(doc=6923)
  0.15 = coord(3/20)
```
Content

"Fast alles wird heute per Computer erfasst. Kaum einer überblickt noch die enormen Datenmengen, die sich in Unternehmen, Universitäten und Verwaltung ansammeln. Allein in den öffentlich zugänglichen Datenbanken der Genforscher fallen pro Woche rund 4,5 Gigabyte an neuer Information an. "Vom potentiellen Wissen in den Datenbanken wird bislang aber oft nur ein Teil genutzt", meint Stefan Wrobel vom Lehrstuhl für Wissensentdeckung und Maschinelles Lernen der Otto-von-Guericke-Universität in Magdeburg. Sein Doktorand Mark-Andre Krogel hat soeben mit einem neuen Verfahren zur Datenbankrecherche in San Francisco einen inoffiziellen Weltmeister-Titel in der Disziplin "Data-Mining" gewonnen. Dieser Daten-Bergbau arbeitet im Unterschied zur einfachen Datenbankabfrage, die sich einfacher statistischer Methoden bedient, zusätzlich mit künstlicher Intelligenz und Visualisierungsverfahren, um Querverbindungen zu finden. "Das erleichtert die Suche nach verborgenen Zusammenhängen im Datenmaterial ganz erheblich", so Wrobel. Die Wirtschaft setzt Data-Mining bereits ein, um das Kundenverhalten zu untersuchen und vorherzusagen. "Stellen sie sich ein Unternehmen mit einer breiten Produktpalette und einem großen Kundenstamm vor", erklärt Wrobel. "Es kann seinen Erfolg maximieren, wenn es Marketing-Post zielgerichtet an seine Kunden verschickt. Wer etwa gerade einen PC gekauft hat, ist womöglich auch an einem Drucker oder Scanner interessiert." In einigen Jahren könnte ein Analysemodul den Manager eines Unternehmens selbständig informieren, wenn ihm etwas Ungewöhnliches aufgefallen ist. Das muss nicht immer positiv für den Kunden sein. Data-Mining ließe sich auch verwenden, um die Lebensdauer von Geschäftsbeziehungen zu prognostizieren. Für Kunden mit geringen Kaufinteressen würden Reklamationen dann längere Bearbeitungszeiten nach sich ziehen. Im konkreten Projekt von Mark-Andre Krogel ging es um die Vorhersage von Protein-Funktionen. Proteine sind Eiweißmoleküle, die fast alle Stoffwechselvorgänge im menschlichen Körper steuern. Sie sind daher die primären Ziele von Wirkstoffen zur Behandlung von Erkrankungen. Das erklärt das große Interesse der Pharmaindustrie. Experimentelle Untersuchungen, die Aufschluss über die Aufgaben der über 100 000 Eiweißmoleküle im menschlichen Körper geben können, sind mit einem hohen Zeitaufwand verbunden. Die Forscher möchten deshalb die Zeit verkürzen, indem sie das vorhandene Datenmaterial mit Hilfe von Data-Mining auswerten. Aus der im Humangenomprojekt bereits entschlüsselten Abfolge der Erbgut-Bausteine lässt sich per Datenbankanalyse die Aneinanderreihung bestimmter Aminosäuren zu einem Protein vorhersagen. Andere Datenbanken wiederum enthalten Informationen, welche Struktur ein Protein mit einer bestimmten vorgegebenen Funktion haben könnte. Aus bereits bekannten Strukturelementen versuchen die Genforscher dann, auf die mögliche Funktion eines bislang noch unbekannten Eiweißmoleküls zu schließen.- Fakten Verschmelzen - Bei diesem theoretischen Ansatz kommt es darauf an, die in Datenbanken enthaltenen Informationen so zu verknüpfen, dass die Ergebnisse mit hoher Wahrscheinlichkeit mit der Realität übereinstimmen. "Im Rahmen des Wettbewerbs erhielten wir Tabellen als Vorgabe, in denen Gene und Chromosomen nach bestimmten Gesichtspunkten klassifiziert waren", erläutert Krogel. Von einigen Genen war bekannt, welche Proteine sie produzieren und welche Aufgabe diese Eiweißmoleküle besitzen. Diese Beispiele dienten dem von Krogel entwickelten Programm dann als Hilfe, für andere Gene vorherzusagen, welche Funktionen die von ihnen erzeugten Proteine haben. "Die Genauigkeit der Vorhersage lag bei den gestellten Aufgaben bei über 90 Prozent", stellt Krogel fest. Allerdings könne man in der Praxis nicht davon ausgehen, dass alle Informationen aus verschiedenen Datenbanken in einem einheitlichen Format vorliegen. Es gebe verschiedene Abfragesprachen der Datenbanken, und die Bezeichnungen von Eiweißmolekülen mit gleicher Aufgabe seien oftmals uneinheitlich. Die Magdeburger Informatiker arbeiten deshalb in der DFG-Forschergruppe "Informationsfusion" an Methoden, um die verschiedenen Datenquellen besser zu erschließen."

Baumgartner, R.: Methoden und Werkzeuge zur Webdatenextraktion (2006) 0.01

0.00793935 = product of:
  0.0793935 = sum of:
    0.020242194 = weight(_text_:und in 5808) [ClassicSimilarity], result of:
      0.020242194 = score(doc=5808,freq=14.0), product of:
        0.044633795 = queryWeight, product of:
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.02013827 = queryNorm
        0.4535172 = fieldWeight in 5808, product of:
          3.7416575 = tf(freq=14.0), with freq of:
            14.0 = termFreq=14.0
          2.216367 = idf(docFreq=13101, maxDocs=44218)
          0.0546875 = fieldNorm(doc=5808)
    0.059151303 = weight(_text_:methoden in 5808) [ClassicSimilarity], result of:
      0.059151303 = score(doc=5808,freq=4.0), product of:
        0.10436003 = queryWeight, product of:
          5.1821747 = idf(docFreq=674, maxDocs=44218)
          0.02013827 = queryNorm
        0.56680036 = fieldWeight in 5808, product of:
          2.0 = tf(freq=4.0), with freq of:
            4.0 = termFreq=4.0
          5.1821747 = idf(docFreq=674, maxDocs=44218)
          0.0546875 = fieldNorm(doc=5808)
  0.1 = coord(2/20)

Abstract: Das World Wide Web kann als die größte uns bekannte "Datenbank" angesehen werden. Leider ist das heutige Web großteils auf die Präsentation für menschliche Benutzerinnen ausgelegt und besteht aus sehr heterogenen Datenbeständen. Überdies fehlen im Web die Möglichkeiten Informationen strukturiert und aus verschiedenen Quellen aggregiert abzufragen. Das heutige Web ist daher für die automatische maschinelle Verarbeitung nicht geeignet. Um Webdaten dennoch effektiv zu nutzen, wurden Sprachen, Methoden und Werkzeuge zur Extraktion und Aggregation dieser Daten entwickelt. Dieser Artikel gibt einen Überblick und eine Kategorisierung von verschiedenen Ansätzen zur Datenextraktion aus dem Web. Einige Beispielszenarien im B2B Datenaustausch, im Business Intelligence Bereich und insbesondere die Generierung von Daten für Semantic Web Ontologien illustrieren die effektive Nutzung dieser Technologien.

Borgman, C.L.; Wofford, M.F.; Golshan, M.S.; Darch, P.T.: Collaborative qualitative research at scale : reflections on 20 years of acquiring global data and making data global (2021) 0.01

0.007878931 = product of:
  0.052526206 = sum of:
    0.017508736 = weight(_text_:software in 239) [ClassicSimilarity], result of:
      0.017508736 = score(doc=239,freq=2.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.21915624 = fieldWeight in 239, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.0390625 = fieldNorm(doc=239)
    0.017508736 = weight(_text_:software in 239) [ClassicSimilarity], result of:
      0.017508736 = score(doc=239,freq=2.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.21915624 = fieldWeight in 239, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.0390625 = fieldNorm(doc=239)
    0.017508736 = weight(_text_:software in 239) [ClassicSimilarity], result of:
      0.017508736 = score(doc=239,freq=2.0), product of:
        0.07989157 = queryWeight, product of:
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.02013827 = queryNorm
        0.21915624 = fieldWeight in 239, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.9671519 = idf(docFreq=2274, maxDocs=44218)
          0.0390625 = fieldNorm(doc=239)
  0.15 = coord(3/20)

Abstract: A 5-year project to study scientific data uses in geography, starting in 1999, evolved into 20 years of research on data practices in sensor networks, environmental sciences, biology, seismology, undersea science, biomedicine, astronomy, and other fields. By emulating the "team science" approaches of the scientists studied, the UCLA Center for Knowledge Infrastructures accumulated a comprehensive collection of qualitative data about how scientists generate, manage, use, and reuse data across domains. Building upon Paul N. Edwards's model of "making global data"-collecting signals via consistent methods, technologies, and policies-to "make data global"-comparing and integrating those data, the research team has managed and exploited these data as a collaborative resource. This article reflects on the social, technical, organizational, economic, and policy challenges the team has encountered in creating new knowledge from data old and new. We reflect on continuity over generations of students and staff, transitions between grants, transfer of legacy data between software tools, research methods, and the role of professional data managers in the social sciences.

Berendt, B.; Krause, B.; Kolbe-Nusser, S.: Intelligent scientific authoring tools : interactive data mining for constructive uses of citation networks (2010) 0.01

0.007716874 = product of:
  0.051445827 = sum of:
    0.017148608 = weight(_text_:23 in 4226) [ClassicSimilarity], result of:
      0.017148608 = score(doc=4226,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.23759183 = fieldWeight in 4226, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.046875 = fieldNorm(doc=4226)
    0.017148608 = weight(_text_:23 in 4226) [ClassicSimilarity], result of:
      0.017148608 = score(doc=4226,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.23759183 = fieldWeight in 4226, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.046875 = fieldNorm(doc=4226)
    0.017148608 = weight(_text_:23 in 4226) [ClassicSimilarity], result of:
      0.017148608 = score(doc=4226,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.23759183 = fieldWeight in 4226, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.046875 = fieldNorm(doc=4226)
  0.15 = coord(3/20)

Date: 23. 1.2011 16:02:43

Frické, M.: Big data and its epistemology (2015) 0.01

0.007716874 = product of:
  0.051445827 = sum of:
    0.017148608 = weight(_text_:23 in 1811) [ClassicSimilarity], result of:
      0.017148608 = score(doc=1811,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.23759183 = fieldWeight in 1811, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.046875 = fieldNorm(doc=1811)
    0.017148608 = weight(_text_:23 in 1811) [ClassicSimilarity], result of:
      0.017148608 = score(doc=1811,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.23759183 = fieldWeight in 1811, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.046875 = fieldNorm(doc=1811)
    0.017148608 = weight(_text_:23 in 1811) [ClassicSimilarity], result of:
      0.017148608 = score(doc=1811,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.23759183 = fieldWeight in 1811, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.046875 = fieldNorm(doc=1811)
  0.15 = coord(3/20)

Date: 20. 3.2015 18:23:25

Teich, E.; Degaetano-Ortlieb, S.; Fankhauser, P.; Kermes, H.; Lapshinova-Koltunski, E.: ¬The linguistic construal of disciplinarity : a data-mining approach using register features (2016) 0.01

0.007716874 = product of:
  0.051445827 = sum of:
    0.017148608 = weight(_text_:23 in 3015) [ClassicSimilarity], result of:
      0.017148608 = score(doc=3015,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.23759183 = fieldWeight in 3015, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.046875 = fieldNorm(doc=3015)
    0.017148608 = weight(_text_:23 in 3015) [ClassicSimilarity], result of:
      0.017148608 = score(doc=3015,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.23759183 = fieldWeight in 3015, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.046875 = fieldNorm(doc=3015)
    0.017148608 = weight(_text_:23 in 3015) [ClassicSimilarity], result of:
      0.017148608 = score(doc=3015,freq=2.0), product of:
        0.07217676 = queryWeight, product of:
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.02013827 = queryNorm
        0.23759183 = fieldWeight in 3015, product of:
          1.4142135 = tf(freq=2.0), with freq of:
            2.0 = termFreq=2.0
          3.5840597 = idf(docFreq=3336, maxDocs=44218)
          0.046875 = fieldNorm(doc=3015)
  0.15 = coord(3/20)

Date: 12. 6.2016 20:23:08

Search (51 results, page 1 of 3)

Authors

Years

Languages

Themes