Próxima parada: ¡Big Data?

1 Próxima parada: ¡Big Data?Inforsalud Marzo ‘14 ...
Author: Sandalio Tadeo
0 downloads 8 Views

1 Próxima parada: ¡Big Data?Inforsalud Marzo ‘14

2 Índice Big Data. Paradigma o Moda Arquitectura de Nueva GeneraciónCasos de Uso Referencias

3 Big Data. Paradigma o ModaMODA. Uso, modo o costumbre que está en boga durante algún tiempo PARADIGMA del Griego Paradeima = Modelo, tipo, Ejemplo © Informática El Corte Inglés |

4 Big Data. Paradigma o Moda“Tenemos por primera vez una economía basada en el recurso clave –Información- que además es renovable, y autogenerativo” Big Data es el “siguiente” recurso natural de los Sistemas de Información Oil which is the fuel for modern economy for centuries. However, Oil in its raw form has little value. It needs to be refined and separated into a large number of consumer products, from petrol and kerosene to asphalt and chemical reagents used to make plastics and pharmaceuticals. It is also used in manufacturing a wide variety of materials. Big Data is just like oil, in it’s raw form it provide no value to enterprise, until it is processed and valuable and actionable business insights are “distilled”. Just like the technology that made available 100 years ago to discover oil and process it to consumable products. Big Data technology is going to transform and revolutionize the way enterprise get and use. 4 4 4

5 Big Data. Paradigma o ModaSistemas Operacionales Big Data. Paradigma o Moda Integración de Datos Información “disponible” no analizada © Informática El Corte Inglés |

6 Big Data. Paradigma o ModaHexa Peta Tera Giga Mega Kilo Datawarehouse mes sem día hora min seg ms Positively impacting the instrumented, interconnected, and intellegent consumer is supported by IBM;s Big Data platform. This platform supports the three V’s like no one else. The concept of Data in Motion is a response to the need to deploy analytics MUCH faster than traditional technology and Data at Rest supports the need for deeper analysis on unstructured or non-traditional data that all communications enterprises are being flooded with. Dato NO estructurado 6

7 Big Data. Paradigma o ModaEl porcentaje de datos disponibles que una empresa puede analizar decrece en relación proporcional a la disponibilidad de los mismos. 0,8 Zb (*) 1 Zb 1,8 Zb estimado 35 Zb Volumen Datos mundiales Datos DISPONIBLES para una organización Today organizations are only tapping in to a small fraction of the data that is available to them The challenge if figuring out how to analyze ALL the data, and find insights in these new and unconventional data types Imagine if you could analyze the 12B TB of tweets being created each day to figure out what people are saying about your products, figure out who the key influencers are within your target demographics. Can you imagine being able to mine this data to identify new market opportunities. What if hospitals could take the thousands of sensor readings collected every hour per patients in ICUs to identify subtle indications that the patient is becoming unwell, days earlier that is allowed by traditional techniques. Imagine if a green energy company could use PBs of weather data along with massive volumes of operational data to optimize asset location and utlization, making these environmentally friends energy sources more cost competitive with traditional sources. Imagine if you could make risk decisions, such as whether or not someone qualifies for a mortgage, in minutes, by analyzing many sources of data, including real-time transactional data, while the client is still on the phone or in the office. Image if law enforcement agencies could analyze audio and video feeds in real-time without human intervention to identify suspicious activity. As these new sources of data continue to grow in volume, variety and velocity, so too does the potential of this data to revolutionize the decision-making processes in every industry. In this slide you can see a graph – it’s not to scale, but you get the point – and this graph shows that the percentage of data available to an enterprise is growing enormously; you can see that at the top bar. And as the amount of data available to an organization grows, the percentage of data that the organization can actually process is decreasing. It’s kind of like we’re getting “dumber” as organizations – in terms of proportion of measurement to the data we are collecting - are understanding less and less of it. +CLICK+ I call the shaded area between these opposite trending lines “The Blind Spot”: it contains signals and noise. This area has got all this data in there, and perhaps it would make sense for us to ingest this into our traditional analytic systems, but we don’t know if that data will yield value or not – it’s a blind spot. We have a hunch that there is value in there, but truly we have no idea what’s in the shaded area. Furthermore, while we know there is value in here, we know it’s not all going to be valuable, so how do we sift through the noise to find the signals? We can start ingesting 10 TB a day of data, ask the CIO for her approval for triple OPEX and CAPEX costs on a hunch? So we have to find a way to find signals from the noise in a cost effective manner. Now if we can leverage some new approach to find the value in the blind spot, at a relatively low cost, if we could tie together things like Big Data social media around our core trusted information that we know about our customers, and drop the stuff that isn’t related to what the business is trying to accomplish, you could really start to monetize that relationship and intents - not just transactions. And that’s the difference, right? How do we monetize intent and relationships? - And that’s a problem domain that includes Big Data. In the previous paragraph I just gave a ubiquitous example, since social media is so obviously tied to Big Data. But you can imagine this dichotomy in any industry. For example, think Oil and Gas (O&G) and the well readings streaming in – and wanting to apply analytics to that with geological data that is unstructured and comes from other sources in various formats and is likely often changing (from an attribute perspective). Harvesting wind energy, traffic patterns, and more. Datos que una organización puede PROCESAR (*) Zb (Zettabyte) = 10 3 Exabyte = 10 6 Petabyte = 10 9 Terabyte

8 Big Data. Paradigma o ModaDatos transaccionales y de aplicación Datos Máquina (M2M) Datos Sociales Contenido Empresarial Volumen Estructurado Throughput Velocidad Semiestructurados Ingestión Variedad Altamente desestructurados Veracidad Variedad Altamente desestructurados Volumen Big data comes from many sources. Its much more than traditional data sources. And it order to capitalize on the breakthrough opportunities we’ve discussed, you definitely need to look beyond traditional data sources. But at the same time, don’t forget that big data comes from those traditional sources too. Transactional data and application data is growing an a significant rate. Although it’s structured, that data is large and it is contained in many different structures. Big data includes machine data – logs, web logs, instrumentation data, network data. Data generated by machines is multiplying quickly, and it contains valuable insights that need to be discovered. Social data also needs to be incorporated. Most social data is really textual data. And the valuable insights remain locked within that text and its many possible meanings. And most of that data isn’t valuable, or has a very short expiry date during which it is valuable. That makes social data very challenging – extracting insight from largely textual content in very little time. And enterprise content must be amalgamated as well. And that data comes in many forms, and also in significant volume.

9 Big Data. Paradigma o ModaVolumen Velocidad Variedad Veracidad Datos en Reposo Datos en Movimiento Datos con múltiples formatos Datos ruidosos Fiabilidad de los datos: desfasados, incom- pletos, conflictivos, irónicos, equivocados, vagos, erróneos Deben procesarse TB-EB Datos en “streaming”, no almacenados, decision necesaria en ms Estructurados, no estructurados, texto, multimedia Examples of all 4 volumes: Volume: log files to Banking Sessions to the behavioural analysis or analysis of the customer's segments to the model production Velocity: Fraud in the environment ATM -> autom. Decision in milisecs Variety: Bsp. Customer's contracts (unstructured), risks in these contracts by legal formulations Veracity: Social Media - Blogs, posting: " Bank x is not good " -> bank x must qualify this information (irony, importance!) Grande App Clásicas Tiempo Real M2M No estructurados Docs Corporativos Calidad Social Media 9

10 Posibilidad de analizar mayores volúmenes de informaciónArquitectura de Nueva Generación Posibilidad de analizar mayores volúmenes de información TRADICIONAL Analiza pequeños conjuntos de datos Información Analizada disponible BIG DATA Analiza TODA la Información Análisis de TODA la Información disponible

11 Menor esfuerzo en la carga y limpieza de datosArquitectura de Nueva Generación Menor esfuerzo en la carga y limpieza de datos TRADICIONAL Cuidadosa limpieza de la información ANTES de analizarla Poca cantidad de información cuidadosamente organizada BIG DATA Análisis de la información como está. Limpieza según se necesite Grandes cantidades de información desordenada

12 Se exploran TODOS los datos y se identifican correlacionesArquitectura de Nueva Generación Los datos y CORRELACIONES aportan nuevo valor a la información TRADICIONAL Comenzamos haciendo hipótesis y probamos contra los datos seleccionados Hipotesis Preguntas Datos Respuestas BIG DATA Se exploran TODOS los datos y se identifican correlaciones Datos Exploración Correlación Conocimiento

13 Los datos se analizan “in motion”, como son generados, en Tiempo RealArquitectura de Nueva Generación Los datos “in motion” ofrecen decisiones en tiempo real TRADICIONAL Se analizan los datos DESPUÉS de que han sido procesados y archivados en un Data Warehouse o Data Mart Repositorio Conocimiento Análisis Datos BIG DATA Los datos se analizan “in motion”, como son generados, en Tiempo Real Datos Conocimiento Análisis

14 Source: Matt Eastwood, IDCArquitectura de Nueva Generación Todo tipo de datos Volumenes extensos Contenidos de valor pero dificiles de extraer Puede ser extremadamente sensible a la velocidad de proceso Big Data es un término ampliamente usado porque la tecnología hace posible analizar TODA la información disponible BIG DATA “Big data describe una nueva generación de tecnologías y arquitecturas diseñadas para la extracción economica de valores a partir de ingentes volumenes de una amplia variedad de datos, habilitando alta velocidad en la captura, descubrimiento y/o analisis.” Source: Matt Eastwood, IDC © Informática El Corte Inglés |

15 Arquitectura de Nueva Generación© Informática El Corte Inglés |

16 Big Data en Sanidad para predecir, prevenir y personalizar2- En sanidad. Para hacer Qué? Big Data en Sanidad para predecir, prevenir y personalizar Mejora en la atención sanitaria Investigación Genómica Operativa Clínica Autocuidado A partir de la secuenciación del genoma humano completo. La inteligencia de toda esa información nos conducirá a un nuevo nivel en sanidad, a la llamada medicina personalizada, saber el tratamiento correcto para el paciente correcto El mundo de la salud genera una ingente cantidad y variedad de ddaattoos s ttaannttoo estructurados como no estructurados (recetas e informes escritos a mano, grabaciones, imágenes, etc). El procesamiento y análisis de todos estos datos puede generar nuevas formas de inteligencia y propiciar una operativa clínica más efectiva y eficaz Estamos asistiendo al aauuggee ddee nnuueevvaass herramientas que nos ayudan a ayudarnos a nosotros mismos. Combinando los sensores y los smartphones con el poder de Big Data podemos prevenir, evitar o mitigar enfermedades. De las áreas más interesantes de aplicación de Big Data es, la propia gestión de la atención sanitaria, bbiieenn ddeesdsdee eell llaaddoo ddee las aseguradoras o desde las administraciones públicas. Las plataformas actuales permiten explorar y medir, en segundos, billones de datos clínicos , de gestión, financieros o de las redes sociales para ofrecer inteligencia de negocio © Informática El Corte Inglés |

17 3- Fuentes de información en BIG DATAAtención Primaria y Especializada Prescripción y dispensación farmacéutica Imagen Medica Imágenes “no radiológicas Actividad de urgencias Historia Clínica Electronica Bancos de DNA Hábitos y estilos de vida Genómica Estilo del vida Morbilidad Mortalidad Registro de enfermedades prevalentes Salud Pública Datos disponibles Datos en curso © Informática El Corte Inglés |

18 4- Nuevos Servicios en Sanidad gracias al Big DataNuevos servicios gracias a BIG DATA Tipología de Servicio Asistencia Adecuada Fuente Información Diseño de modelo predictivos en base a características comunes de la población que ha presentado determinadas patologías, identificando población de riesgo e informándolos a tiempo para la prevención y detección precoz (Identificación y comunicación a personas con riesgo de un ICTUS) Historia Clínica Electrónica A Investigación de las patologías mas prevalentes a partir de datos genéticos (Identificación de las principales característica de patologías como cáncer y esquizofrenia.) Asistencia Adecuada Innovación Adecuada Banco de ADN B C Identificación de les tendencias de las patologías que sufre la población en base a la estudios estadísticas epidemiológicos. Asistencia Adecuada Innovación Adecuada Salud Pública Historia Clínica Electrónica D Determinación de la seguridad y eficacia de nuevos tratamientos en base la definición y análisis de indicadores clave Proveedor Adecuado Salud Pública E Optimización de estudios clínicos y que empresas farmacéuticas pagarían para la obtención de información de valor para el desarrollo de nuevos fármacos Innovación Adecuada Historia Clínica Electrónica Salud Pública F Óptima gestión de los recursos en base a su disponibilidad geográfica y según la actividades asistenciales en las áreas de salud Proveedor Adecuado Historia Clínica Electrónica Salud Pública G Identificación de los principales factores de riesgo en la aparición de patologías y fomentar hábitos saludables en la población, aconsejando de esta manera que la población esté informada y se corresponsabilice en la prevención Estilo de vida Adecuado Valor Adecuado Estilo de vida H Investigación de los motivos de las readmisiones para ahorrar costes en los centros sanitarios Asistencia Adecuada Historia Clínica Electrónica Salud Pública © Informática El Corte Inglés |

19 National Health Service (NHS).5- Algunas iniciativas National Health Service (NHS). La iniciativa Prescribing Analitycs en su primer trabajo ha analizado las recetas de dos estatinas (fármacos utilizados para disminuir el colesterol en sangre) prescritas en el National Health Service (NHS). Las estatinas originales se venden con su marca comercial y tienen una alternativa en forma de medicamento genérico mucho más barato. Lo que hace el estudio es representar en un mapa de Inglaterra las porcentajes de estas estatinas por marcas y genéricos, obteniendo una representación visual Departament de Salut (AIAQS) Se quiere implementar un modelo que permita facilitar las grandes cantidades de datos que se generan continuamente en el sistema de salud de Catalunya a todos los agentes que intervienen o tienen capacidad para mejorar la salud de la población: ciudadanía, profesionales, investigadores, gestores, organizaciones sanitarias, gobiernos, empresas, centros de investigación, centros tecnológicos, empresas de tecnología sanitaria, etc. de la prescripción muy impactante y que permite identificar posibles ahorros (potenciales) en el gasto producido por estos medicamentos. © Informática El Corte Inglés |

20 Infografía en TicBeat © Informática El Corte Inglés |

21 Detección de Síntomas en Tiempo RealEl Instituto de la Universidad de Ontario detecta los síntomas de neonatos con anterioridad Ejecuta analítica en tiempo real utilizando datos fisiológicos de los neonatos Correlaciona datos continuamente de monitores médicos para detectar cambios sútiles y alertar al personal médico antes El sistema avisa a los cuidadores de posibles complicaciones Beneficios: Ayuda a detectar condiciones de amenaza hasta 24 horas antes Reducción de mortandad infantil y mejora de los cuidados de los pacientes Client name: University of Ontario Institute of Technology Subtitle: Leveraging key data to provide proactive patient care The need: Today, patients are routinely connected to equipment that continuously monitors vital signs such as blood pressure, heart rate and temperature. The equipment issues an alert when any vital sign goes out of the normal range, prompting hospital staff to take action immediately, but many life-threatening conditions do not reach critical level right away. Often, signs that something is wrong begin to appear long before the situation becomes serious, and even a skilled and experienced nurse or physician might not be able to spot and interpret these trends in time to avoid serious complications. One example of such a hard-to-detect problem is nosocomial infection, which is contracted at the hospital and is life threatening to fragile patients such as premature infants. The indication is a pulse that is within acceptable limits, but not varying as it should. So, while the information needed to detect the infection is present, the indication is very subtle; rather than being a single warning sign, it is a trend over time that can be difficult to spot. The solution/benefit: With a shared interest in providing better patient care, Dr. Carolyn McGregor, Canada Research Chair in Health Informatics at the University of Ontario Institute of Technology (UOIT), and Dr. Andrew James, staff neonatologist at The Hospital for Sick Children (SickKids) in Toronto, partnered to find a way to make better use of the information produced by monitoring devices. Dr. McGregor visited researchers at the IBM T.J. Watson Research Center’s Industry Solutions Lab (ISL), who were extending a new stream-computing platform to support healthcare analytics. A three-way collaboration was established, with each group bringing a unique perspective—the hospital focus on patient care, the university’s ideas for using the data stream, and IBM providing the advanced analysis software and information technology expertise needed to turn this vision into reality. The result was Project Artemis, a highly flexible platform that aims to help physicians make better, faster decisions regarding patient care for a wide range of conditions. The earliest iteration of the project is focused on early detection of nosocomial infection by watching for reduced heart rate variability along with other indications. For safety reasons, in this development phase the information is being collected in parallel with established clinical practice and is not being made available to clinicians. The early indications of its efficacy are very promising. Project Artemis is based on IBM InfoSphere Streams. The IBM DB2 relational database provides the data management required to support future retrospective analysis of the collected data. Detección de Síntomas en Tiempo Real 21

22 Análisis de Imágenes para DiagnósisBureau Salud Asiático reduce errores de diagnóstico Necesidad El servicio telemédico de diagnóstico por imágenes tiene como objetivo aumentar la salud rural Automaticamente mueve y analiza grandes collecciones de imágnes buscando anomalías y enfermedades Hace posible que radiólogos y patólogos analicen 1000s imágenes de pacientes cada día Mejoras esperadas: Reducción en errores de diagnóstico Resultados mejorados aprovechando el tratamiento médico de casos similares We’ve been working with a health bureau in China helping them develop a centralized medical imaging diagnostics solution. Challenges: An estimated 80 percent of healthcare data is medical imaging data, in particular radiology imaging. Telemedicine is becoming an effective approach in developing countries to enhance the medical image collaboration among large hospital and rural hospitals who often lack dedicate radiologists. There is a great need for analytics that data and quickly shift through these large collections looking for relevant information, such as anomalies or disease - confirming imaging. Physicians routinely examine 1000s of X-Rays and 10,000s of MRIs per day during the course of diagnosing patients. Reviewing image after image results in significant eye fatigue making it easy to overlook certain anomalies besides the one under investigation. With limited time to process the daily load of patients, human observation errors occur. Such undetected ailments in the long run drive up the cost of healthcare. Developing a medical imaging analytics system to automatically detect anomalies has been a significant challenge, not only for running disease-specific algorithms, but also for the huge data volume process. Parallel processing is needed to mine these large imaging data sets, which can range anywhere from 10 of gigabytes, to terabytes or even petabytes. Solution: IBM’s Hadoop system (InfoSphere BigInsights) was chosen because of its proven high performance enterprise class big data platform and the ability to run very compute intensive medical imaging algorithms. Benefit: The medical imaging diagnostics platform is expected to significantly improve patient healthcare by allowing physicians to exploit the experience of other physicians in treating similar cases, and inferring the prognosis and the outcome of treatments. Allows physicians to see consensus opinions as well as differing alternatives, helping reduce the uncertainty associated with diagnosis. In the long run, these capabilities will lower diagnostic errors and improve the quality of care. Análisis de Imágenes para Diagnósis 22

23 © Informática El Corte Inglés | www.ieci.es

24 © Informática El Corte Inglés | www.ieci.es

25 © Informática El Corte Inglés | www.ieci.es